Customer Io to ClickHouse: How to Move Your Data
Move Customer.io into ClickHouse with Airbyte. Why ten requests per second sets your backfill, and which streams actually justify a column store.

Moving Customer.io into ClickHouse gives you fast analysis over messaging activity across long periods. Customer.io reports on campaigns competently and becomes slow once somebody asks how delivery and engagement have moved across two years of sends.
This guide covers the managed path with Airbyte. Two things shape the build: a fixed rate limit makes sync duration something you plan around rather than tune, and only some of the streams are large enough to justify this destination at all.
Customer Io to ClickHouse at a glance:
Why move data from Customer Io to ClickHouse?
Two situations account for most of these pipelines.
The first is interactive analysis of messaging at scale. Delivery, open and click behaviour across millions of sends is heavy aggregation, and the difference between a query taking a second and a minute changes how often anybody investigates a drop in engagement.
The second is joining messaging to revenue and product usage. If your interest is mainly in the campaign structures rather than the activity, those are small and a warehouse suits them better, which is what Customer Io to BigQuery covers.
What do you need before you start?
Four things, and the first is easy to get wrong because Customer.io has several:
An App API key, specifically. Not the credentials your application uses to send messages, which belong to a different API entirely. The Customer.io source documentation sets out which key the connector expects.
Agreement that behavioural events are out of scope. The events your application pushes into Customer.io are served by a different host from the one this connector uses, so they are not part of this pipeline.
A realistic estimate of your activity volume. Ten requests per second is a fixed ceiling, so the size of your messaging history is the length of your first sync and there is no setting that changes it.
A ClickHouse database and a sorting key per table. Almost every query against messaging data bounds a period, so the sorting key usually begins with the date.
If your database restricts inbound traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.
How do you build a Customer Io to ClickHouse pipeline in Airbyte?
Step 1: Work out how long the backfill will take
Take your volume of activity and messages, set it against ten requests per second, and accept the answer. This limit is a property of the host the connector uses rather than something your plan or configuration influences, so the arithmetic is the whole forecast. Doing it now converts an unexplained multi-day sync into a planned one, which is the difference between a patient stakeholder and an anxious one.
Step 2: Configure the Customer.io source
Click Sources in the left navigation, then New Source, and select Customer.io, following adding a source. Supply the App API key, then select streams deliberately, since campaigns, newsletters and segments are small while activities and messages are where your volume lives.
Step 3: Configure the ClickHouse destination
Click Destinations, then New Destination, and select ClickHouse, following adding a destination. Supply the host, port, database and credentials. Records land in typed columns over the native protocol, and the types settled at table creation are the ones you keep, so inspect them after the first sync.
Step 4: Create the connection and split the schedules
Click Connections, then New connection, select your streams and a sync mode. Consider separate connections for the small configuration streams and the large activity ones, since a campaign definition changing weekly does not need to drag a full activity sync behind it against a fixed rate limit.
Then build views that unnest campaign structures, because those arrive with actions, templates and segment identifiers inside them.
Why is ten per second the whole story?
Because it belongs to the API host rather than to your account. Customer.io runs several APIs, this connector uses one of them, and that host permits ten requests per second. There is no tier that lifts it and no connector setting that works around it, which makes it unusually simple to reason about and impossible to negotiate with.
For an organisation sending at volume, that turns the first sync into a scheduling question. A history of millions of messages will take as long as the arithmetic says, and the interface will report everything as healthy throughout, because it is. Restarting discards progress and changes nothing about the ceiling.
Plan around it rather than against it. Start earlier than the deadline that prompted the project, set a start date covering the history you genuinely need rather than everything ever sent, and separate small streams onto their own connection so routine updates are not queued behind a large one. Steady state is comfortable; it is the first pass that costs.
Which streams justify a column store?
The activity ones, and really only those. Campaigns, newsletters, segments and sender identities describe how your messaging is configured, and there are hundreds of them rather than millions. Nothing about that data needs a column store, and a small table anywhere would serve it perfectly well.
Activities and messages are the opposite, growing with every send and every interaction, and they are the reason to be here. Aggregating open rates by segment across two years, or comparing delivery across campaigns, is exactly the shape of query a column store answers in a second and a row store labours over.
So treat the two groups differently. Sort the large tables by date first, since almost every question bounds a period, and let the small configuration tables be whatever is convenient. Unnest campaign structures into views rather than querying inside them repeatedly, because actions, templates and segment identifier arrays are awkward to reach into and the definitions rarely change enough to justify doing it on every query.
Frequently asked questions
Can I get the events my app sends to Customer.io?
Not through this connector. Those are served by a different Customer.io host, so behavioural events you push in are outside this pipeline entirely.
Can I speed up the sync?
Not meaningfully. Ten requests per second is a property of the host rather than your plan, so narrow the history and the streams instead.
Which key does the connector need?
An App API key, which is distinct from the credentials your application uses to send messages. Supplying the wrong one is a common first stumble.
Why are campaign fields awkward to query?
They arrive nested, carrying actions, templates and arrays of segment identifiers. Unnest them once in a view rather than reaching inside on every query.
Can I do this without writing code?
The pipeline, yes. The views unnesting campaign structures and the sorting keys on your large tables are SQL, and both are what makes this fast rather than merely populated.
Get your Customer Io data into ClickHouse
Use an App API key and set expectations that behavioural events are not part of this. Do the arithmetic on ten requests per second before promising a date, because that ceiling is fixed and the backfill takes what it takes. Then separate the tiny configuration streams from the large activity ones, sort the big tables by date, and unnest campaign structures into views rather than reaching into them on every query.
Airbyte's connector catalog includes 600+ pre-built connectors, so messaging activity can be analysed beside the outcomes it produced. For mobile attribution into the same destination, see Appsflyer to ClickHouse, and for storefront data into the same destination, Woocommerce to ClickHouse.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
