Create a dedicated user.
Create anairbyte_user rather than reusing an existing one. Set async_insert = 0 on it, which the connector requires for connection checks and syncs to work correctly when async inserts are enabled on the server.How it works
Create a user, grant it, point Airbyte at your host. There is no bucket to provision and no storage integration to configure.
airbyte_user rather than reusing an existing one. Set async_insert = 0 on it, which the connector requires for connection checks and syncs to work correctly when async inserts are enabled on the server.CREATE, ALTER, TRUNCATE, INSERT, SELECT, CREATE DATABASE, CREATE TABLE, and DROP TABLE on the target database. Repeat the same grants for any custom namespace.Not a philosophy. A set of specific, verifiable guarantees about the infrastructure your data runs on.
Why Airbyte
A dedicated user and network access are enough. Columns stay typed. Writes stay batched.
The connector’s requirements are a ClickHouse instance, network access, and a user. Object storage is not among them.
Direct Load writes data into columns matching your source schema rather than storing everything as JSON in raw tables.
Batches are capped at 70 MB, and record window size is tunable if you need to.
ClickHouse has excellent first-party ingestion for the shapes it was designed around, particularly Kafka and object storage. The gap is everything that starts as a SaaS API.
| Method | Best for | Limits | Use Airbyte instead when |
|---|---|---|---|
| Kafka table engine | Streaming events already on a Kafka topic | Kafka only. You own the producer and the topic | Your data is not on a topic and putting it there is the actual work |
| s3() and url() table functions | Files already sitting in object storage | You own extraction, file formatting, and scheduling | Your data starts in an API or a database |
| ClickPipes | Managed ingestion on ClickHouse Cloud | ClickHouse Cloud only, and a bounded source list | You are self-hosted, or your source is not covered |
| PeerDB | Postgres change data capture into ClickHouse | Postgres only. No SaaS connectors | You need Postgres CDC and SaaS sources in the same cluster |
| clickhouse-client bulk insert | One-off backfills and manual loads | Not a pipeline. No scheduling, no state, no schema handling | You need this to keep running after you stop watching it |
Breadth into one cluster. Postgres CDC plus Salesforce, HubSpot, and Zendesk landing in the same ClickHouse database on the same schedule, with schema drift handled, is not something the first-party options cover together.
For a pure Kafka-to-ClickHouse pipeline, the Kafka table engine is simpler.
All 700+ of them. Most ClickHouse users arrive from Postgres.
Databases
SaaS applications
Airbyte writes typed columns directly, and delegates deduplication to ClickHouse’s own engine rather than running MERGE queries against your cluster.
| Airbyte type | ClickHouse type |
|---|---|
| String | String |
| Integer | Int64 |
| Decimal | Decimal(38, 9) |
| Boolean | Bool |
| Timestamp | DateTime64(3) |
| Object | JSON if you enable JSON in the connector configuration, otherwise String |
| Array | String |
| Union | String |
Column added.
Airbyte adds the column to the destination table automatically.
Type changed.
Airbyte modifies the column type as needed.
Column removed.
ClickHouse is the exception. Airbyte drops the column and its historical data from the destination table, because ReplacingMergeTree deduplicates by deleting and re-inserting records rather than using MERGE, and that table recreation requires dropping obsolete columns.
The numbers
0M
pipelines synced daily
18K+
companies
199%
ROI, certified by Forrester
Deployment
Your warehouse credentials and your data do not have to leave your environment. Airbyte does not retain customer data on its servers.
Fully managed hosting on the Standard and Pro plans. Fastest path to a running pipeline.
Hybrid deployment, with Airbyte’s data plane running in your cloud, VPC, or on-prem environment. You get 700+ connectors and a fully managed experience, while your data never touches our servers.
Open source, your infrastructure, your rules.
compliance
SOC 2 Type II certified, GDPR and HIPAA support, with tools to help you meet internal and external regulatory requirements.
24/7 support and 99.9% availability backed by contractual SLAs and priority response times for mission-critical workloads.
Move data into your warehouse/lakehouse of choice, VPC, on-prem, or into Airbyte's Context Store.
Didn’t find your answer?
Please don’t hesitate to reach out.
What is the best way to load data into ClickHouse?
Does Airbyte need a staging bucket for ClickHouse?
Why do my queries return duplicate rows?
What happens when a column is removed from my source?
Which cursor column should I use for incremental syncs?
What permissions does Airbyte need in ClickHouse?
Can I run this against self-hosted ClickHouse?
Airbyte clones tables instead of swapping them, so the grants you configured survive every sync. Removed source columns are kept, and oversized values are logged rather than fatal.
Airbyte stages files in Google Cloud Storage, then loads them into BigQuery with native load jobs. Warehouse compute stays on Google Cloud, and Airbyte never sits in the query path.
Airbyte stages Avro files inside Unity Catalog Volumes, not an external bucket, so your data stays under your catalog’s governance from the first byte to the final table.