Gitlab to Snowflake: How to Move Your Data

Move GitLab into Snowflake with Airbyte. Why blank scope fields take everything, why group projects cap at 100, and what a primary key change costs.

Summarize with AI:

Moving GitLab into Snowflake gives engineering leaders the delivery history their tooling never assembles in one place. Merge request cycle time, pipeline reliability and how those have changed across teams are trend questions, and GitLab answers them one project at a time if at all.

This guide covers the managed path with Airbyte. Two things shape the build: the scope of what you sync is decided by fields that are optional and unhelpful when left blank, and a past version changed primary keys in a way that makes upgrades more than a version bump.

Gitlab to Snowflake at a glance:

CapabilitySupportedWhat it means for this pipeline
APIGitLab v4Works with gitlab.com and self-hosted instances
Scope fieldsBoth optionalLeft blank, you get every group the token can see
Projects per groupCapped at 100A reported limitation, and it truncates silently
Token scoperead_apiRead-only across groups, projects and registries
Version 4.0.0Changed primary keysSix streams need a schema refresh and a reset

Why move data from Gitlab to Snowflake?

Two situations account for most of these pipelines.

The first is measuring delivery across an organisation rather than a project. Cycle time and review latency are derived from timestamps spread over merge requests, commits and pipelines, and comparing them between teams means having all of it in one governed place.

The second is joining engineering activity to incidents, releases or cost. If several systems need to react to GitLab events rather than analyse them, that is a different shape and Gitlab to Kafka suits it better than a warehouse will.

What do you need before you start?

Four things, and the second is the one people leave empty by accident:

A personal access token with the read_api scope, or OAuth. That scope grants read-only access across groups, projects and the registries. The GitLab source documentation covers both routes.

An explicit list of groups or projects. Both fields are optional, and leaving both blank means the connector takes every group your token can reach and syncs their projects, which is almost never what anybody intended.

Your API URL if you self-host. The connector works with both gitlab.com and self-hosted instances, and the default assumes the former.

Snowflake objects and a role. A warehouse, database, schema and a role that can create tables. Engineering activity data is modest by warehouse standards, so sizing is rarely the interesting question.

If your instance or Snowflake account restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow lists before you begin.

How do you build a Gitlab to Snowflake pipeline in Airbyte?

Step 1: Name your groups or projects explicitly

Decide which parts of your GitLab estate the analysis covers, and type them into the groups or projects field rather than leaving both empty. An empty configuration is not a neutral default here: it means everything the token can reach, which on an organisation account is a great deal more data, a much longer sync and a dataset whose boundary nobody can describe afterwards.

Step 2: Configure the GitLab source

Click Sources in the left navigation, then New Source, and select GitLab, following adding a source. Authenticate with OAuth or your token, set the API URL if you self-host, supply a start date, and list your groups or projects. Then select streams, which cover merge requests, commits, pipelines, issues and the surrounding metadata.

Step 3: Configure the Snowflake destination

Click Destinations, then New Destination, and select Snowflake, following adding a destination. Supply the account identifier, warehouse, database, schema and role. Merge request and pipeline records carry nested structures, and Snowflake holds those natively, so nothing needs flattening on the way in.

Step 4: Create the connection and count your projects

Click Connections, then New connection, select your streams and a sync mode. After the first sync, compare the number of projects that arrived against the number you expect, for the reason covered below. Daily is ample, since nobody needs a merge request modelled within minutes.

Then build views, because the landing tables describe GitLab's objects rather than your delivery process.

What does the connector actually cover?

Whatever you named, or everything if you named nothing. The groups and projects fields are both optional, and the documented behaviour when both are blank is that the connector retrieves all groups accessible to the token and syncs their projects. That is a reasonable default for a small account and an unpleasant surprise on a large one.

The sharper issue is a reported limitation on how projects within a group are retrieved. Because the connector reads a group's projects from the group endpoint rather than paginating the projects endpoint, a group containing more than a hundred projects yields only a hundred. No error appears, the sync succeeds, and the missing projects are simply absent.

That combination is worth planning around. If any of your groups is large, list projects explicitly rather than relying on group expansion, and in either case compare the project count in your warehouse against what GitLab reports after the first sync. Check the connector's changelog too, since this is the sort of limitation that gets fixed and the version you are running determines whether it applies to you.

What happens when primary keys change?

Rather more than a version bump, which is why version 4.0.0 is worth knowing about even if you are starting fresh today. It changed the primary key on six streams, covering group and project members, group and project labels, branches and tags, and required a schema refresh and a reset of the affected streams.

A primary key change matters in a warehouse because it is what deduplication is built on. Resetting a stream means re-landing its data, and anything downstream that assumed the old key, including views joining on it, needs revisiting rather than merely rerunning. That is a morning of work when planned and an afternoon of confusion when not.

The general lesson is worth applying beyond this one release. Have everything downstream read views rather than the landing tables directly, so a future key or schema change is one view definition to update instead of a hunt through everybody's dashboards. And read the connector's breaking change notes before upgrading, because they name the affected streams precisely and the alternative is inferring it from a broken report.

Frequently asked questions

What happens if I leave groups and projects blank?

The connector takes every group your token can access and syncs their projects. Name your scope explicitly unless you genuinely want the whole estate.

Why are some projects missing?

A reported limitation caps projects retrieved through a group at one hundred, without erroring. List projects explicitly for large groups and check your connector version.

Which token scope do I need?

read_api, which grants read-only access to the API including groups, projects and the registries. OAuth is the alternative.

Does this work with a self-hosted GitLab?

Yes. The connector uses GitLab API v4 and works with both gitlab.com and self-hosted instances, with the API URL field pointing at yours.

Can I do this without writing code?

The pipeline, yes. Deriving cycle time and review latency from timestamps across several streams is modelling work, and it is where the value appears.

Get your Gitlab data into Snowflake

Name your groups or projects rather than leaving both fields blank, since an empty configuration means the whole estate. Count the projects that arrive against what you expect, because retrieval through a group is capped at a hundred and says nothing about it. Read the breaking change notes before upgrading, given that version 4.0.0 changed primary keys on six streams. Then have reporting read views, so the next change is one definition to update.

Airbyte's connector catalog includes 600+ pre-built connectors, so engineering activity can be measured beside everything it affects. For the same source into a warehouse on another platform, see Gitlab to BigQuery, and for the other major source control platform into the same destination, Github to Snowflake.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.