Github to ClickHouse: How to Move Your Data
Move GitHub into ClickHouse with Airbyte. Why most incremental streams still read everything, and why repeated records make FINAL the default.

Moving GitHub into ClickHouse gives you fast aggregation over the streams that actually pile up. Comments, reviews and workflow runs accumulate constantly on an active organisation, and asking how build duration has drifted across a year is exactly what a column store is for.
This guide covers the managed path with Airbyte. Two things shape the build: most incremental streams are less incremental than the word suggests, and that read behaviour decides how you query the resulting tables.
Github to ClickHouse at a glance:
Why move data from Github to ClickHouse?
Two situations account for most of these pipelines.
The first is interactive analysis of the high-volume streams. Comment activity, review throughput and workflow duration across a year are heavy aggregations, and the difference between a second and a minute decides whether anybody investigates a trend or assumes one.
The second is powering a dashboard engineering teams actually watch. If you want delivery metrics modelled and joined to commercial data instead, Github to Databricks suits that better and handles nested records natively.
What do you need before you start?
Four things, and the second is the one that surprises people:
A GitHub token and a named repository list. Leaving the field blank takes everything the token can see, which on an organisation account is far more than you meant. The GitHub source documentation describes how each stream behaves.
Realistic expectations of incremental sync. Only four streams read just the new records. Most others read everything and emit only what changed, which affects your API budget rather than your table sizes.
A ClickHouse database and a sorting key per table. Nearly every question about engineering activity bounds a period, so the sorting key usually begins with a date.
A plan for deduplication. Because records arrive more than once here as a matter of routine rather than as an occasional accident.
If your database restricts inbound traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.
How do you build a Github to ClickHouse pipeline in Airbyte?
Step 1: Find out how your chosen streams actually read
Check each stream you want against the connector's notes on incremental behaviour, because the label hides three different arrangements. Four streams genuinely read only new records. Most read everything and output only changes. The workflow streams revisit the recent past on every run. Knowing which is which explains both your sync duration and the duplicates you will see later.
Step 2: Configure the GitHub source
Click Sources in the left navigation, then New Source, and select GitHub, following adding a source. Supply the token, repository list and a start date. A very distant start date on a large stream can produce errors rather than records, so be moderate about history on the heavy streams.
Step 3: Configure the ClickHouse destination
Click Destinations, then New Destination, and select ClickHouse, following adding a destination. Supply the host, port, database and credentials. Records land in typed columns over the native protocol, and the types settled at table creation are the ones you keep.
Step 4: Create the connection and spread the load
Click Connections, then New connection, select your streams and a sync mode. Splitting heavy streams into their own connection with a slower schedule is the recommended way to live within the rate limits, and it keeps the streams you check daily arriving on time.
Then build the views everybody reports from, applying FINAL where a figure leaves the building.
Why do incremental streams still read everything?
Because GitHub's API does not let the connector ask every endpoint for changes since a date. Four streams are pure incremental, reading only new records and outputting only new records: comments, commits, issues and review comments. Those behave the way the word implies and cost you little.
The larger group is different. Those streams read all records and output only the new ones, which means your tables stay tidy while your API budget is spent as though you were doing a full refresh. The connector's own documentation flags this as something to bear in mind precisely because it affects call limits rather than anything you can see in the data.
The workflow streams are a third case, revisiting roughly the past thirty days on each run so that runs which completed after the last sync are caught. That is sensible behaviour and it means recent records arrive repeatedly by design. All three patterns push in the same direction: fewer streams per connection, heavy ones on their own slower schedule, and a token budget you have actually thought about.
What does that mean once the data lands?
That repeated records are the steady state rather than an edge case. A workflow run caught while in progress and again once finished produces two rows, and the thirty day revisit means recent history keeps being re-sent. Both are correct behaviour and both leave you with several versions of the same object.
In a column store that matters because deduplication happens during background merges rather than on write, so for a period both versions sit in the table. Counting workflow runs directly will overstate them, and averaging duration will include records that had none because the run had not finished when it was read.
So make FINAL the default in any view anybody reports from, rather than a remedy you apply after noticing a strange number. Sort these tables by the date the object was created or started, since every question bounds a period, and keep status and conclusion columns typed for low cardinality because they take a handful of repeated values across millions of rows.
Frequently asked questions
Why is an incremental sync so slow?
Most incremental streams read all records and output only the new ones. Only four are pure incremental, so the saving is in your tables rather than in API calls.
Why do recent workflow runs keep reappearing?
Those streams revisit roughly the past thirty days so that runs finishing after a sync are captured. Report through views applying FINAL.
How do I stay inside the rate limits?
Use incremental where it helps, sync less often, and split heavy streams into separate connections. Remember REST and GraphQL are metered differently.
My workflow failure rate looks wrong.
You are probably counting in-progress and completed versions of the same run. Apply FINAL and filter to terminal statuses for anything about outcomes.
Can I do this without writing code?
The pipeline, yes. The views applying FINAL and the sorting keys are SQL, and they are what makes this accurate rather than merely fast.
Get your Github data into ClickHouse
Check how each chosen stream reads before scheduling anything, because only four are pure incremental and most spend your API budget as though they were doing a full refresh. Split the heavy ones onto their own slower connection. Then treat repeated records as normal rather than as a fault, make FINAL the default in reporting views, and sort by date with low cardinality status columns.
Airbyte's connector catalog includes 600+ pre-built connectors, so engineering activity can be analysed at whatever speed the questions demand. For the same source into a search engine, see Github to Elasticsearch, and for the other major source control platform into the same destination, Gitlab to ClickHouse.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
