Github to Teradata: How to Move Your Data
Move GitHub into Teradata with Airbyte. Why an API source leaves nowhere to flatten nested records, and why each stream is a modelling exercise.

Moving GitHub into Teradata brings engineering activity into the environment where an organisation already models everything else. Where finance, operations and customer data live in governed subject areas, delivery metrics sitting in a separate tool is an exception somebody has to explain.
This guide covers the managed path with Airbyte. Two things shape the build: an API source gives you nowhere to reshape data before it lands, and in a strictly modelled environment every stream you select is a piece of modelling work.
Github to Teradata at a glance:
Why move data from Github to Teradata?
Two situations account for most of these pipelines.
The first is bringing engineering data to where the analysts are. An organisation with a mature Teradata practice has reporting conventions, access models and people who know the platform, and adding delivery metrics there beats asking everyone to learn a second environment.
The second is joining delivery activity to cost or incident data that already lives there. The caveat is structural: GitHub records nest heavily, and if you want to work with that shape rather than flatten it, Github to Databricks reads it natively and saves you the exercise.
What do you need before you start?
Four things, and the first is not the safe option it appears to be:
A chosen SSL mode. Encryption is off by default and two of the six modes permit an unencrypted connection, so leaving the default is a decision rather than an omission. The Teradata destination documentation lists the modes and logon mechanisms.
A GitHub token and a named repository list. Leaving the repository field blank takes everything the token can see, which on an organisation account is far more modelling work than anybody signed up for. The GitHub source documentation covers the options.
A schema name your organisation recognises. Tables land in a default schema called airbyte_td unless you specify otherwise, which is not a name that belongs in a governed environment.
A short list of the metrics you actually want. Because here the stream list is the project plan, for reasons the second half of this guide explains.
If your Teradata system restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.
How do you build a Github to Teradata pipeline in Airbyte?
Step 1: Set the SSL mode deliberately
Choose an encryption mode before anything else, because the default does not encrypt and two of the available modes will happily connect without it. In an organisation running Teradata there is almost certainly a policy on this, and the person who will ask about it is easier to satisfy before the pipeline exists than after somebody notices during a review.
Step 2: Configure the GitHub source
Click Sources in the left navigation, then New Source, and select GitHub, following adding a source. Supply the token, your repository list and a start date, then select streams sparingly. Remember that REST counts requests while GraphQL calculates points, so a wide selection loads both budgets differently.
Step 3: Configure the Teradata destination
Click Destinations, then New Destination, and select Teradata, following adding a destination. Supply the host, credentials, logon mechanism and SSL mode, and name the schema rather than accepting the default. Expect nested structures to arrive in whatever form the destination can accommodate rather than as modelled columns.
Step 4: Create the connection, then start modelling
Click Connections, then New connection, select your streams and a sync mode. Use incremental where offered, since engineering history only grows, and daily is ample. The landing tables are the beginning of the work here rather than the end of it.
Check numeric and text column definitions while you are there, since a governed environment tends to notice loose typing faster than a lake does.
Why is there nowhere to flatten these records?
Because the source is an API rather than a database. When you replicate from a warehouse into Teradata you can create a view upstream that unnests everything and sync that instead, putting the reshaping where the nested types are understood. GitHub offers no such place: the API returns what it returns.
That matters because pull request records nest generously. An author object, an array of labels, a list of requested reviewers and references to branches and repositories all sit inside one record, and a strictly relational destination has no native shape for any of it. The reshaping is unavoidable and it happens after landing rather than before.
So plan the modelling as part of the project rather than as a tidy-up afterwards. Decide which nested fields become proper columns, which become their own tables with a foreign key back to the pull request, and which you simply discard because nobody will query them. That last category is usually larger than people expect and deciding it deliberately is what keeps the model comprehensible.
Which streams justify the modelling effort?
Very few, and that is the useful discipline this destination imposes. In a lakehouse an extra stream costs storage and nothing else, so taking everything is reasonable. Here each stream is a table somebody has to type properly, name conventionally and integrate into an existing subject area, which is real work per stream.
Work backwards from the metrics. Review latency and cycle time need pull requests, reviews and commits, which is three streams and a handful of derived durations. Almost everything else in the GitHub catalogue, reactions and comment streams especially, feeds no question anybody has actually asked and would cost the same modelling effort as the streams that do.
The rate limits happen to agree with that conclusion. The heavier streams tend to spend points on the GraphQL budget rather than requests on the REST one, so a narrow selection chosen for modelling reasons also syncs faster. When the source and the destination push you towards the same answer, it is usually the right one.
Frequently asked questions
Is the connection encrypted by default?
No. SSL is off by default and two of the six modes permit unencrypted connections, so choose a mode deliberately rather than accepting what is offered.
Can I flatten nested records before they arrive?
Not with an API source, since there is no upstream view to create. Plan the reshaping as work you do in Teradata after the data lands.
Where do my tables end up?
In a schema named airbyte_td unless you specify another, which is worth doing so they sit where your organisation expects to find things.
Should I take every available stream?
No. Each stream is a modelling exercise in this environment, so select the few that feed your delivery metrics and leave the rest.
Can I do this without writing code?
The pipeline, yes. Reshaping nested records into a model and deriving delivery metrics from timestamps across streams is SQL, and it is most of the project.
Get your Github data into Teradata
Set the SSL mode rather than accepting a default that does not encrypt, and name a schema your organisation recognises. Then select streams from your metrics list rather than from the catalogue, because each one is a modelling exercise here and nested pull request records have to be reshaped after landing rather than before. Conveniently, the narrow selection that suits the modelling also suits GitHub's rate limits.
Airbyte's connector catalog includes 600+ pre-built connectors, so engineering activity can be modelled beside everything else the business measures. For the same source into another relational database, see Github to MySQL, and for an enterprise database into the same destination, Oracle Database to Teradata.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
