ClickHouse to MongoDB: How to Move Your Data

Move ClickHouse into MongoDB with Airbyte. Why you should ship aggregates rather than tables, and why derived data makes full refresh the sensible choice.

Summarize with AI:

Moving ClickHouse into MongoDB is a serving pattern rather than a migration. ClickHouse answers analytical questions across enormous tables and is not what you want behind an application fetching one customer's summary a thousand times a second, which is exactly what a document store is for.

This guide covers the managed path with Airbyte. Two things shape the build: what you move should be computed results rather than raw rows, and because those results are derived, several rules that normally apply to replication stop applying here.

ClickHouse to MongoDB at a glance:

CapabilitySupportedWhat it means for this pipeline
Sync modesFull refresh and incrementalCursor based, using the JDBC driver
Incremental deletesListed as coming soonSo deletions do not reach the destination today
Logical replicationListed as coming soonThere is no change capture option to choose
Cursor columnsRequired for incrementalA table without one can only full refresh
IndexesYour responsibilityThe pipeline creates collections and never indexes

Why move data from ClickHouse to MongoDB?

One situation genuinely suits this, and it is worth being precise about it.

The good case is serving computed results. You have aggregations that take real work to produce, an application needs them by key in milliseconds, and running that aggregation on every request would be both slow and wasteful. Computing once and serving many times is the pattern, and a document store is a good place for the serving half.

The poor case is moving raw analytical tables, since MongoDB aggregates worse than the system you are leaving and the volumes are unkind. If your application needs relational access to the results rather than document fetches, ClickHouse to PostgreSQL serves the same purpose in a shape more applications already expect.

What do you need before you start?

Four things, and the first decides whether this pipeline is small or unmanageable:

An aggregate to move, rather than a table. Build a view or table in ClickHouse holding one row per thing your application fetches, and sync that. Moving raw rows and aggregating on the other side inverts the strengths of both systems.

ClickHouse connection details and a cursor if you want incremental. The connector uses the ClickHouse JDBC driver and supports SSL and SSH tunnelling. The ClickHouse source documentation sets out the features and their status.

A document design matching the fetch. One document per key the application asks for, carrying everything that request needs, which is the only reason to be using a document store at all.

An index plan. The pipeline creates collections and never creates indexes, so without them every lookup is a scan and the speed you came for is gone.

If either system restricts inbound traffic by IP, add the Airbyte Cloud IP addresses to the allow lists before you begin.

How do you build a ClickHouse to MongoDB pipeline in Airbyte?

Step 1: Build the aggregate in ClickHouse first

Write down exactly what the application fetches and by which key, then create something in ClickHouse producing one row per that key. This is the step that makes the whole arrangement sensible: the heavy work happens where the column store is strong, and what travels is small and already shaped. Skip it and you are moving analytical volumes into a system that will neither store them comfortably nor query them well.

Step 2: Configure the ClickHouse source

Click Sources in the left navigation, then New Source, and select ClickHouse, following adding a source. Supply the host, port, database and credentials, selecting your aggregate rather than the underlying tables. Namespaces are enabled by default, and the connector does not alter your schema.

Step 3: Configure the MongoDB destination

Click Destinations, then New Destination, and select MongoDB, following adding a destination. Supply the connection string, database and credentials. Documents arrive with metadata fields the destination adds, so application code should request fields by name rather than assuming a document holds only your own.

Step 4: Create the connection, then build the indexes

Click Connections, then New connection, select your stream and a sync mode. Once the first load completes, index the field your application queries by before pointing anything at the collection, because an unindexed lookup is a scan and defeats the purpose of the exercise.

Match the schedule to how stale the application can tolerate the figures being, which for most aggregates is hours rather than minutes.

What should you actually move?

The answer, not the working. The temptation is to replicate tables and let the application aggregate, which puts the heavy lifting in the system least suited to it and moves far more data than necessary. The productive version computes in ClickHouse and ships a result small enough that the pipeline is almost incidental.

That reframes what the connector's limitations mean. Incremental sync needs a cursor column, and both incremental deletes and logical replication are listed as coming soon rather than available, which would be a serious constraint if you were mirroring a live table. For a recomputed aggregate it barely matters, because you are replacing a derived set rather than tracking changes to a source of truth.

So design the aggregate around the request. One document per customer, per account or per whatever key the application holds, containing everything that screen needs in a single fetch. If two screens need different shapes, that is two collections rather than one clever document, since storage is cheap and a second fetch is what you were trying to avoid.

Why does derived data change the rules?

Because a recomputed aggregate has no history to protect. Most replication advice exists to preserve fidelity with a source of truth: capture deletes, avoid overwriting, keep what you cannot recover. None of that applies to a table you can rebuild at any moment from data that still exists in ClickHouse.

That makes full refresh a reasonable choice rather than a last resort, and it dissolves the deletes problem entirely: replacing the whole set removes anything that no longer belongs, without the connector needing to know a deletion occurred. Where an application is already reading the collection, load into a new one and repoint rather than overwriting in place, so a failed load never leaves consumers with nothing.

What does still apply is the indexing, because nothing about derived data makes a collection scan fast. Create your indexes as part of the load procedure rather than as a one-off, since a repoint pattern means new collections regularly and an unindexed one will be noticed immediately by whoever is waiting on the page.

Frequently asked questions

Should I move raw tables or aggregates?

Aggregates. Compute in ClickHouse and ship one row per key the application fetches, since MongoDB aggregates worse than the system you are leaving.

Will deletions propagate?

Not today, since incremental deletes and logical replication are both listed as coming soon. For a recomputed aggregate, replacing the whole set handles it anyway.

Why can a table not sync incrementally?

It has no column usable as a cursor. Tables are only offered for incremental sync when at least one column can serve that purpose.

Is overwriting safe here?

Safer than usual, because the data is derived and rebuildable. Still load into a new collection and repoint where an application is reading, so a failed load never empties it.

Can I do this without writing code?

The pipeline, yes. The aggregate in ClickHouse and the indexes in MongoDB are both yours, and they are what the whole arrangement depends on.

Get your ClickHouse data into MongoDB

Build the aggregate in ClickHouse before configuring anything, because shipping computed results rather than raw rows is what makes this pairing sensible. Shape one document per key the application fetches. Then use the fact that the data is derived: full refresh is coherent, missing delete support stops mattering, and a load-then-repoint pattern protects your consumers. Just never forget the indexes, which the pipeline will not create for you.

Airbyte's connector catalog includes 600+ pre-built connectors, so analytical results can reach the applications that display them. For the same source into another serving store, see ClickHouse to Convex, and for a relational source into the same destination, MySQL to MongoDB.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.