Klaviyo to BigQuery: How to Move Your Data

Move Klaviyo into BigQuery with Airbyte. Why the optional fields have expensive defaults, and what disabling predictive analytics quietly removes.

Summarize with AI:

Moving Klaviyo into BigQuery lets you judge email and SMS against what they actually produced. Klaviyo reports on opens and clicks competently and cannot tell you whether a flow generated customers who stayed, because the revenue and retention data lives elsewhere.

This guide covers the managed path with Airbyte. Two things shape the build: several optional configuration fields are expensive when left blank, and one reliability setting removes a field that downstream work may depend on.

Klaviyo to BigQuery at a glance:

CapabilitySupportedWhat it means for this pipeline
CredentialPrivate API keyWith scopes matching each stream you want
Start dateOptionalLeft blank, everything is replicated
Conversion metric IDsOptionalLeft blank, reports fetch every metric and run slowly
Lookback window0 days by defaultAnd applies only to the events detailed stream
Predictive analyticsCan be disabledImproves Profiles reliability, removes the field entirely

Why move data from Klaviyo to BigQuery?

Two situations account for most of these pipelines.

The first is attribution you control. Klaviyo will tell you what it believes a flow earned, and checking that against your own order and refund data is a different exercise that can only happen where both datasets sit together.

The second is keeping history beyond what the interface makes convenient, and joining profiles to product usage or support contacts. If your questions are heavy aggregations over very large event volumes and query speed is the point, a column store suits that better than a warehouse will.

What do you need before you start?

Four things, and three of them are fields you can technically skip:

A private API key with the right scopes. Each stream needs the scope for its endpoint, so check which ones your selection requires rather than creating a key and hoping. The Klaviyo source documentation explains how to find them.

A start date you actually chose. The field is optional and leaving it empty replicates everything, which on an account sending for years is a considerably longer first sync than anybody planned.

Your conversion metric identifiers, if you want the reports. Without them the campaign and flow report streams fetch every metric, which the documentation itself notes can be slow because of rate limits.

A BigQuery dataset in the right location. Location is fixed at creation and BigQuery will not join across locations, so put this where your order and product data already live.

If your network restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow list before you begin.

How do you build a Klaviyo to BigQuery pipeline in Airbyte?

Step 1: Fill in the optional fields deliberately

Go through the optional settings before configuring anything else and decide each one on purpose. Start date, conversion metric identifiers and lookback window are all skippable, and the defaults for the first two are the expensive choices rather than the neutral ones. Five minutes here is the difference between a sync that finishes overnight and one somebody investigates on Thursday.

Step 2: Configure the Klaviyo source

Click Sources in the left navigation, then New Source, and select Klaviyo, following adding a source. Supply the private API key, your start date, any conversion metric identifiers and a lookback window. Then select streams, noting that the report streams are the ones most affected by the metric setting.

Step 3: Configure the BigQuery destination

Click Destinations, then New Destination, and select BigQuery, following adding a destination. Supply the project identifier, dataset and service account credentials. Events are the volume here, so a busy sending account is one of the cases where Cloud Storage staging is worth considering over batched inserts.

Step 4: Create the connection and set a lookback if you need one

Click Connections, then New connection, select your streams and a sync mode. The lookback window defaults to zero days and applies to the detailed events stream, so raise it if late-arriving events matter to you and leave it alone if they do not.

Then have reporting read views that filter on the extraction timestamp partition, since BigQuery bills on bytes scanned and event tables grow quickly.

Why are the optional fields the expensive ones?

Because skipping them means taking everything. An empty start date replicates all data rather than none, which is the safer default from the connector's point of view and the slower one from yours. On an account with several years of sending, that is the difference between a first sync measured in hours and one measured in days.

The conversion metric setting works the same way and is less obvious. Leaving it empty makes the campaign and flow report streams fetch reports for every metric in your account, and the documentation warns plainly that this can be slow because of rate limits. Most organisations care about two or three conversion metrics, so naming them turns a laborious stream into a quick one.

The lookback window is the exception, in that its default of zero is usually fine and its scope is narrower than people assume: it applies to the detailed events stream rather than to everything. If your analysis depends on events that arrive late, raise it deliberately for that stream. If it does not, leaving it at zero avoids re-reading data you already have.

What do you give up to make Profiles reliable?

A field, and possibly something built on it. The Profiles stream can hit transient API errors under heavy load, and the connector offers a setting to disable fetching predictive analytics which improves the success rate of those syncs. It is a sensible remedy and it is not free.

The documentation is explicit about the cost: records on the Profiles stream will no longer contain the predictive analytics field, and anything depending on that field stops working. In a warehouse that means a column which was populated last month is empty this month, with no error anywhere, which is exactly the kind of change that reaches a dashboard before anybody notices.

So treat it as a decision with downstream consequences rather than a troubleshooting toggle. Check whether any view, model or dashboard reads predictive analytics before switching it off, and if you do switch it off, say so somewhere findable. Reliability is usually worth more than a field nobody queries, and the failure mode when somebody does query it is silent enough to deserve a note.

Frequently asked questions

Why is my first sync taking so long?

Probably an empty start date, which replicates everything, or empty conversion metric identifiers, which makes report streams fetch every metric. Both are optional fields with expensive defaults.

The Profiles stream keeps failing.

It can hit transient errors under load. Disabling predictive analytics improves the success rate, at the cost of that field disappearing from records entirely.

Does the lookback window apply to every stream?

No, it applies to the detailed events stream and defaults to zero days. Raise it only if late-arriving events matter to your analysis.

A stream returns nothing.

Check the scopes on your private API key, since each stream needs the scope for its endpoint and a key created quickly often misses one.

Can I do this without writing code?

The pipeline, yes. The views joining Klaviyo activity to your own revenue data, and filtering on the partitioning column, are SQL worth writing early.

Get your Klaviyo data into BigQuery

Fill in the optional fields on purpose, because an empty start date takes everything and empty conversion metric identifiers make the report streams fetch every metric slowly. Check your key's scopes against the streams you selected. Then decide about predictive analytics deliberately rather than as a fix, since disabling it removes the field and breaks anything reading it without raising an error anybody will see.

Airbyte's connector catalog includes 600+ pre-built connectors, so messaging can be judged against the revenue it claims. For another messaging platform into the same destination, see Customer Io to BigQuery, and for product analytics into the same destination, Amplitude to BigQuery.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.