Azureblobstorage to Amazon S3 with AWS Glue: How to Move Your Data

Move Azure Blob Storage into S3 with AWS Glue using Airbyte. Why the delivery method decides the project, and why your schema comes from the files.

Summarize with AI:

Moving Azure Blob Storage into Amazon S3 with AWS Glue usually happens during a platform move or when analysis lives in one cloud and the files that feed it live in another. Files accumulated in Azure become Iceberg tables an AWS estate can query.

This guide covers the managed path with Airbyte. Two things shape the build: a delivery method choice that decides whether you are moving files or making tables, and the fact that your schema comes from the files themselves.

Azureblobstorage to Amazon S3 with AWS Glue at a glance:

CapabilitySupportedWhat it means for this pipeline
Delivery methodA choiceParse files into records, or move them as they are
SchemaInferred from filesSo the files you have decide the columns you get
AuthenticationKey or service principalThe second can be scoped to the container you need
ScopeOne containerSeveral containers means several sources
MaintenanceYours to scheduleCompaction and snapshot expiry do not happen alone

Why move data from Azureblobstorage to Amazon S3 with AWS Glue?

Two situations account for most of these pipelines.

The first is making files queryable rather than merely present. Turning exports and logs into Iceberg tables registered in a catalogue means Athena, Spark and Redshift Spectrum can read them without anybody writing a parser, which is a different outcome from copying files across.

The second is consolidating during a cloud migration, where analysis has moved to AWS and the producers have not. If you want this data in a warehouse instead, Azure Blob Storage to Snowflake asks far less of you operationally than a lake does.

What do you need before you start?

Four things, and the first decides what kind of project this is:

A decision about the delivery method. The connector can parse your files into records or move them as they are, and those produce very different results. The Azure Blob Storage source documentation covers the options and the file formats.

Credentials, ideally a service principal. A storage account key works, and client credentials let you grant only the container read and blob read permissions this connector needs, which is the better arrangement where your Azure administrators allow it.

A container per source. The configuration names one container, so files spread across several need several sources and a decision about whether they become one table or many.

An S3 bucket, a Glue catalog and a maintenance schedule. Keep namespace and table names alphanumeric with underscores, since Glue rewrites anything else, and plan compaction from the start.

If your storage account restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build an Azureblobstorage to Amazon S3 with AWS Glue pipeline in Airbyte?

Step 1: Choose between files and tables

Decide whether you want the contents of your files as queryable records or the files themselves sitting in S3, because the delivery method setting is where that choice is made and everything downstream follows from it. If the answer is that you simply want the blobs copied, a storage transfer tool may serve you better than a pipeline. If you want Iceberg tables in a catalogue, this is the right instrument.

Step 2: Configure the Azure Blob Storage source

Click Sources in the left navigation, then New Source, and select Azure Blob Storage, following adding a source. Supply the account name, container and credentials, then define your streams by naming a file format and a path pattern. Being specific with that pattern is what stops unrelated files joining your table.

Step 3: Configure the S3 Data Lake destination

Click Destinations, then New Destination, and select the S3 Data Lake, following adding a destination. Choose AWS Glue as the catalog and supply the bucket, region and credentials. Name things in plain letters, digits and underscores so the catalog reads the way you intended.

Step 4: Create the connection and inspect the inferred schema

Click Connections, then New connection, select your streams and a sync mode. Look at the columns and types that were discovered before anybody builds on them, then schedule compaction and snapshot expiry from the first week rather than the first complaint.

Check cross-cloud transfer costs too, since moving a large historical container out of Azure is a bill somebody should see coming.

Are you moving files or making tables?

The delivery method decides, and the two answers serve different projects. Parsing records reads inside your files and produces rows, which is what turns a container of exports into Iceberg tables somebody can query with SQL. Copying raw files moves the objects themselves, preserving them exactly as they are.

That second option is worth being honest about, because if all you need is the same files in a different cloud, a dedicated storage transfer tool is usually faster, cheaper and less to operate than a data pipeline. Choosing raw copy here makes sense when it is one step of a larger flow you are already running through Airbyte rather than the whole job.

For most people reading this, parsing is the point. It is also where the work sits, because a table is only as good as the path pattern and format configuration that produced it. One stream per logical dataset, with a pattern narrow enough to exclude the odd manifest or archive somebody dropped in the container, is what keeps the resulting tables sensible.

What decides the columns you end up with?

The files, which is a different arrangement from a database source where the schema is declared. The connector works out columns and types from the files it finds, so your table reflects what was in the container at discovery rather than a contract anybody wrote down, and the quality of that inference depends heavily on the format.

Parquet and Avro carry their own types and behave well. CSV does not, since everything in it is text until something decides otherwise, and a column of identifiers with leading zeroes or a date in an unusual format is exactly where that decision goes wrong. Check the inferred types before building on them rather than after somebody reports a mangled reference.

The related risk is drift. A producer adding a column to files written next month gives you files that disagree with the ones discovery saw, and while Iceberg can evolve a table's schema, files already written keep the layout they were written with. So agree with whoever produces these files that the shape is now an interface, and re-check the schema when they tell you it changed.

Frequently asked questions

Should I copy raw files or parse records?

Parse records if you want queryable Iceberg tables. If you only want the same files in S3, a storage transfer tool is usually simpler than a pipeline.

Why are my CSV types wrong?

CSV carries no type information, so they were inferred. Check identifiers with leading zeroes and unusual date formats first, and configure the format explicitly where you can.

Can one source cover several containers?

No. The configuration names one container, so several containers need several sources and a decision about whether they land as one table or many.

Unexpected files ended up in my table.

The path pattern is too broad. Narrow it so manifests, archives and anything else living in the container do not match the stream.

Can I do this without writing code?

The pipeline, yes. Compaction and snapshot expiry are jobs you schedule, and a lake without them degrades quietly.

Get your Azureblobstorage data into Amazon S3 with AWS Glue

Settle the delivery method first, because parsing records and copying raw files are different projects and only one of them needs a pipeline. Use a service principal scoped to the container where you can. Then inspect the inferred schema before building on it, treat CSV types with suspicion, and agree with your file producers that the shape is now an interface. Schedule compaction from the first week.

Airbyte's connector catalog includes 600+ pre-built connectors, so files in one cloud can become tables in another. For a relational source into the same destination, see MySQL to Amazon S3 with AWS Glue, and for the same source into a warehouse, Azure Blob Storage to BigQuery.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.