Gitlab to Amazon S3 with AWS Glue: How to Move Your Data

Move GitLab into S3 with AWS Glue using Airbyte. Why a silently truncated archive is worse than a broken dashboard, and what a connector upgrade costs.

Summarize with AI:

Moving GitLab into Amazon S3 with AWS Glue builds a durable record of engineering activity in an open format. Pipelines, commits and merge requests accumulate constantly, and Iceberg tables in a catalogue keep them cheaply while staying readable by several engines.

This guide covers the managed path with Airbyte. Two things shape the build: an archive that quietly omits projects is worse than a dashboard that does, and a connector upgrade can cost you more here than in most destinations.

Gitlab to Amazon S3 with AWS Glue at a glance:

CapabilitySupportedWhat it means for this pipeline
Scope fieldsBoth optionalLeft blank, you get every group the token can see
Projects per groupCapped at 100A reported limitation, and it truncates silently
Version 4.0.0Changed primary keysOn six streams, needing a refresh and a reset
Table maintenanceYours to scheduleCompaction and snapshot expiry do not happen alone
NamingAlphanumericGlue rewrites anything else in namespaces and tables

Why move data from Gitlab to Amazon S3 with AWS Glue?

Two situations account for most of these pipelines.

The first is long-term retention at a sensible price. Engineering history is worth keeping for years and consulted occasionally, so storage cost matters more than query latency, and an open format keeps it readable by whatever you use next.

The second is feeding several engines from one copy. If what you want is interactive speed over build history rather than a durable archive, Gitlab to ClickHouse answers those aggregations considerably faster.

What do you need before you start?

Four things, and the first decides whether your archive means anything:

An explicit list of groups or projects. Both fields are optional and blank means everything the token can reach, which is rarely what anybody intended. The GitLab source documentation covers the options.

A token with the read_api scope, and your API URL if self-hosted. The connector works with gitlab.com and self-hosted instances alike, and the default assumes the former.

An S3 bucket, a Glue catalog and a maintenance schedule. Keep namespace and table names alphanumeric with underscores, since Glue rewrites anything else, and plan compaction from the start.

A position on connector upgrades. Because a version change here has occasionally meant resetting streams, and a reset against a lake table is not the small thing it is elsewhere.

If your instance restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build a Gitlab to Amazon S3 with AWS Glue pipeline in Airbyte?

Step 1: Establish what the archive is meant to cover

Write down which groups and projects this record is supposed to include, then check that the pipeline actually delivers them. An archive is consulted rarely and trusted absolutely, so a gap nobody noticed is a considerably worse problem here than in a dashboard somebody looks at weekly and would question.

Step 2: Configure the GitLab source

Click Sources in the left navigation, then New Source, and select GitLab, following adding a source. Authenticate, set the API URL if you self-host, and name projects explicitly where a group holds more than a hundred, since group expansion is capped.

Step 3: Configure the S3 Data Lake destination

Click Destinations, then New Destination, and select the S3 Data Lake, following adding a destination. Choose AWS Glue as the catalog and supply the bucket, region and credentials, naming things in plain letters, digits and underscores.

Step 4: Create the connection and count what arrived

Click Connections, then New connection, select your streams and a sync mode. Then compare the project count in your tables against what GitLab reports, and schedule compaction and snapshot expiry from the first week rather than the first complaint.

Record the covered scope beside the tables, since nothing in the data describes what was meant to be there.

What is your archive actually covering?

Possibly less than you configured, because two behaviours work against each other. The groups and projects fields are optional, so an empty configuration reaches for everything the token can see, and separately a reported limitation caps projects retrieved through a group at one hundred without raising an error.

On a large organisation those combine badly. Somebody leaves the fields blank expecting complete coverage, the expansion truncates quietly, and the resulting tables look entirely healthy while missing projects nobody thought to check for. There is no error, no warning and no gap in the data that reveals itself.

That matters more in an archive than anywhere else. A dashboard missing a team gets questioned within a fortnight by the team; a table consulted once a year for a compliance question or a retrospective is trusted without anybody re-verifying it. So name projects explicitly for large groups, count what arrived, and keep a written statement of intended scope beside the catalogue.

What does a connector upgrade cost here?

Occasionally a rebuild, which is a bigger event against a lake table than against a warehouse. Version 4.0.0 of this connector changed the primary keys on several streams, including group and project members, labels, branches and tags, and required a schema refresh and a stream reset for the affected ones.

A reset means the table is rewritten rather than amended, and on an Iceberg table holding years of history that is a real operation with a real cost. It also interacts with your maintenance: newly written data arrives as fresh files needing compaction, and old snapshots hang around until expiry removes them.

So treat upgrades as planned work rather than housekeeping. Read the changelog before accepting a major version, do it when somebody can watch, and check afterwards that the rewritten tables hold what they held before. The open format helps here, since your data remains readable throughout, but nothing about it makes a reset free.

Frequently asked questions

Why are some projects missing?

A reported limitation caps projects retrieved through a group at one hundred, without erroring. List projects explicitly for large groups and check your connector version.

What happens if I leave the scope fields blank?

You get every group the token can see, which is rarely intended and interacts badly with the project cap on a large organisation.

Do connector upgrades affect my tables?

They can. Version 4.0.0 changed primary keys on several streams and required a schema refresh and reset, which rewrites the affected tables.

Why are my table names different from what I typed?

Glue rewrites characters outside letters, digits and underscores, so name namespaces and tables plainly if you want them to read as intended.

Can I do this without writing code?

The pipeline, yes. Compaction and snapshot expiry are jobs you schedule, and a lake without them degrades quietly.

Get your Gitlab data into Amazon S3 with AWS Glue

Name your groups and projects explicitly and count what arrived, because the scope fields default to everything and group expansion truncates silently, and an archive nobody re-checks is exactly where that goes unnoticed. Keep names plain for Glue. Then treat connector upgrades as planned work, since a primary key change has meant resetting streams and a reset rewrites a lake table.

Airbyte's connector catalog includes 600+ pre-built connectors, so engineering history can be kept in a format that outlives the tools reading it. For the same source into a warehouse, see Gitlab to Snowflake, and for a relational source into the same destination, MySQL to Amazon S3 with AWS Glue.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.