AWS CloudTrail to BigQuery: How to Move Your Data
Move AWS CloudTrail into BigQuery with Airbyte. Why dataset location is a residency decision you cannot undo, and how to keep an audit archive affordable.

Moving AWS CloudTrail into BigQuery keeps cloud activity long after CloudTrail's lookup forgets it. About ninety days covers an incident response and not a compliance review, so the archive you build is the only thing that answers a question about last year.
This guide covers the managed path with Airbyte. Two things shape the build: where the dataset lives is a governance decision rather than a convenience, and an archive that grows forever is queried at a price.
AWS CloudTrail to BigQuery at a glance:
Why move data from AWS CloudTrail to BigQuery?
Two situations account for most of these pipelines.
The first is retention for audit, since ninety days is short of any reasonable compliance requirement and a warehouse holds years cheaply while keeping the records queryable with SQL your existing tools already speak.
The second is joining cloud activity to cost or identity data you keep elsewhere. If your work is heavy transformation over deeply nested event records, AWS CloudTrail to Databricks handles that shape more comfortably.
What do you need before you start?
Four things, and the first may need somebody outside your team:
A decision about where this data may live. Audit records from a regulated workload often carry residency expectations, and BigQuery dataset location is fixed at creation, so this is worth confirming rather than assuming.
A list of your accounts and regions. Events are recorded per region and looked up per account, so coverage is a multiplication. The AWS CloudTrail source documentation covers the fields.
Credentials with CloudTrail read access. An access key and secret plus the region name, set deliberately since the connector defaults to us-east-1 and a wrong value quietly produces an unexpected dataset.
A view on who will query this and how often. An audit archive is consulted rarely and scanned expensively, which is a combination worth designing around rather than discovering.
If your network restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow list before you begin.
How do you build an AWS CloudTrail to BigQuery pipeline in Airbyte?
Step 1: Settle the dataset location with somebody accountable
Decide where this dataset lives before creating it, because the location cannot be changed afterwards and audit data is the kind most likely to have a rule attached. Moving it later means creating a new dataset and reloading, and with a ninety-day source window anything older than that cannot be reloaded at all. Ask the question once, early, and record the answer.
Step 2: Configure the AWS CloudTrail source
Click Sources in the left navigation, then New Source, and select AWS CloudTrail, following adding a source. Supply the access key, secret key, region and start date, repeating per account and region. Name each source after its account and region, since a list of identical names becomes unmanageable quickly.
Step 3: Configure the BigQuery destination
Click Destinations, then New Destination, and select BigQuery, following adding a destination. Supply the project identifier, the dataset you located deliberately, and service account credentials. A busy estate produces enough events that Cloud Storage staging is worth considering.
Step 4: Create the connections and build the views
Click Connections, then New connection for each source, selecting the stream and a sync mode. Use incremental and schedule comfortably inside ninety days. Then write the views investigators will use, with partition filters built in, before anybody queries the tables directly.
This connector is community maintained rather than certified, so run it for a fortnight before an auditor depends on it.
Where is this audit data allowed to live?
A question this pairing raises more sharply than most, because you are moving records about one cloud provider's infrastructure into another's. CloudTrail events describe who did what in your AWS estate, including principals, source addresses and the resources touched, and that is exactly the material a regulator has opinions about.
The awkward part is that BigQuery's dataset location is fixed when the dataset is created and cannot be altered. Events from a European workload landing in a dataset somebody created in the United States by accident is not a configuration mistake you fix with a setting, it is a reload into a new dataset, and a reload only reaches back as far as the source window allows.
So map your AWS regions against the dataset locations you intend to use, and check that pairing with whoever owns compliance before creating anything. If different regions have different requirements, that means separate datasets rather than one convenient home, and BigQuery will not join across them, which is a design consequence worth knowing at the start rather than the middle.
How do you keep a growing archive affordable?
By filtering the partition, because this is the awkward shape for a warehouse that bills on bytes scanned: a table nobody queries for months, then queried urgently by somebody who does not know its structure. Tables arrive partitioned on the extraction timestamp, and a query ignoring that column reads every event you have ever collected.
That pattern is predictable enough to design for. An investigation starts with a time window and a principal or a resource, so views taking those as parameters and filtering the partition first give investigators the right shape without requiring them to know how the table is laid out. The alternative is somebody scanning three years of events to find one afternoon.
Union the accounts and regions in those views too, with columns naming both, since an investigator should not have to know which of eleven tables holds the answer. And record what each view covers, because an audit dataset whose scope is undocumented is difficult to rely on when somebody asks whether the record is complete.
Frequently asked questions
Can I change the dataset location later?
No. It is fixed at creation, so moving means a new dataset and a reload that can only reach back as far as CloudTrail's ninety-day window.
Why are my queries expensive?
Probably not filtering the extraction timestamp partitioning column, so each query scans the whole archive. Build that filter into the views investigators use.
Can one source cover several regions?
No. Events are recorded per region and looked up per account, so you need a source for each combination and a union view to bring them together.
Can I backfill older activity?
Not through this connector, since the lookup reaches back about ninety days. For a longer record, start collecting now or use a trail or event data store.
Can I do this without writing code?
The pipelines, yes, though there are several. The views that union regions and filter the partition are SQL and they are what makes the archive usable.
Get your AWS CloudTrail data into BigQuery
Settle the dataset location with whoever owns compliance before creating anything, because it cannot be changed and a reload cannot reach past ninety days. Count your accounts against your regions, since that is how many sources you are building. Then write views that union them, filter the partitioning column and take a time window, so an investigator finds the answer without scanning three years of events.
Airbyte's connector catalog includes 600+ pre-built connectors, so cloud activity can be retained long after the platform forgets it. For the same source into a lakehouse, see AWS CloudTrail to Databricks, and for the same source into a search engine, AWS CloudTrail to Elasticsearch.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
