MySQL to Amazon S3 with AWS Glue: How to Move Your Data
Move MySQL into S3 with AWS Glue using Airbyte. Why binlog deletes keep a lake honest, and why partitioning must be decided before the first load.

Moving MySQL into Amazon S3 with AWS Glue puts application data into an open format several engines can read. Years of transactional history are worth keeping and rarely queried daily, which is an expensive combination in a warehouse and a natural one for Iceberg tables on object storage.
This guide covers the managed path with Airbyte. Two things shape the build: change capture delivers deletes, which is unusually valuable for a lake, and the partitioning you choose at creation is the decision the table's performance rests on.
MySQL to Amazon S3 with AWS Glue at a glance:
Why move data from MySQL to Amazon S3 with AWS Glue?
Two situations account for most of these pipelines.
The first is retention at a sensible price. Application history is worth keeping indefinitely and consulted occasionally, so storage cost matters more than query latency, and Iceberg on S3 holds it cheaply while staying readable by Athena, Spark and Redshift Spectrum without another copy.
The second is feeding transformation work with more than one engine. The counterargument is operational, since a lake means catalogs and table maintenance to look after. If you want interactive reporting with none of that, MySQL to Snowflake asks considerably less of you.
What do you need before you start?
Four things, and the second has a default that will catch you out:
A user with the three required grants. SELECT, REPLICATION CLIENT and REPLICATION SLAVE together allow change capture from the binlog. The MySQL source documentation covers the grants and server settings.
Binlog retention set deliberately. On RDS it defaults to zero, and Airbyte recommends 168 hours. Without that, any pause leaves nothing to resume from and recovery means a full resnapshot.
An S3 bucket and a Glue catalog. The pipeline writes data files to one and registers tables in the other. Keep namespace and table names alphanumeric with underscores, since Glue rewrites anything else.
A partitioning decision per table. Chosen from how the data will be read rather than from how MySQL organises it, and settled before the first load rather than after.
If your database restricts inbound traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.
How do you build a MySQL to Amazon S3 with AWS Glue pipeline in Airbyte?
Step 1: Set binlog retention before anything else
On RDS the retention parameter defaults to zero, which means the binlog is discarded immediately and a pipeline that stops for any reason has nothing to resume from. Set it to 168 hours so a weekend outage is recoverable. This is one parameter and it is the difference between a broken sync you restart and a full resnapshot of every table you were careful about.
Step 2: Configure the MySQL source
Click Sources in the left navigation, then New Source, and select MySQL, following adding a source. Supply the host, port, database and credentials, and choose change capture as the replication method. Check your storage engines while you are here, since a snapshot can lock a MyISAM table on a live system.
Step 3: Configure the S3 Data Lake destination
Click Destinations, then New Destination, and select the S3 Data Lake, following adding a destination. Choose AWS Glue as the catalog and supply the bucket, region and credentials. Name things in plain letters, digits and underscores so the catalog reads as you intended.
Step 4: Create the connection and schedule maintenance
Click Connections, then New connection, select your streams and a sync mode. Then schedule compaction and snapshot expiry from the first week rather than the first complaint, and alert on failure, since a pipeline down longer than your retention cannot resume.
Frequent syncs are generous here, since every run writes files and the maintenance burden follows how often you write rather than how much.
Why do deletes matter more in a lake?
Because a lake is the copy people trust when the source no longer has the answer. Binlog change capture delivers deletions alongside inserts and updates, which means your Iceberg tables can stay honest about what the application actually holds. Cursor-based reading never sees a removal, so the copy drifts further from the source every year it runs.
That accuracy has a mechanical cost worth understanding. Iceberg has no update in place, so a delete arriving through change capture is written as a delete file rather than removing anything, and a high-churn table produces those steadily. The table stays correct and the file count climbs, which is precisely what compaction exists to resolve.
So treat the two together. Change capture is worth the grants and the retention parameter because it keeps the lake truthful, and compaction is worth scheduling because that truthfulness is what generates the files. A team that enables one and neglects the other ends up with an accurate table nobody enjoys querying.
How should you partition a table fed from MySQL?
By how it will be read, which is rarely how MySQL organises it. An application table is arranged around the primary key for transactional access, and the queries people will run against the lake copy are almost always bounded by time instead. Partitioning by a date column is therefore the usual answer, and copying the source's organisation across is the usual mistake.
Iceberg handles this more gracefully than older formats, since partitioning is hidden and queries do not need to know about it, and the specification can evolve later if your access patterns change. What evolution does not do is rewrite history: files already written keep the layout they were written with, so a query spanning the change reads across both arrangements.
Which is why it is worth a conversation before the first load rather than after. Ask what periods people will query and how often, partition at a granularity that produces sensibly sized files rather than thousands of tiny ones, and remember that a very high-churn table receiving change capture is already generating delete files, so an over-granular partition scheme compounds a problem you will be compacting anyway.
Frequently asked questions
Why can the pipeline not resume after a pause?
Binlog retention is probably zero, which is the RDS default. Set it to 168 hours, and expect a full resnapshot to recover this time.
Do deletions reach the lake?
With binlog change capture, yes, which is one of the strongest arguments for using it. Cursor-based reading sees only inserts and updates.
Can I change partitioning later?
The specification can evolve, but files already written keep their original layout, so a query spanning the change reads across both. Decide deliberately up front.
Could the first sync affect my application?
If any selected table uses MyISAM, the snapshot can lock it. Check your storage engines and either convert, schedule carefully or exclude those tables.
Can I do this without writing code?
The pipeline, yes, though the grants and retention parameter are database administration. Compaction jobs are yours, and a lake without them degrades quietly.
Get your MySQL data into Amazon S3 with AWS Glue
Set binlog retention to 168 hours before anything else, because the RDS default of zero turns any pause into a resnapshot. Use change capture, since deletions reaching the lake are what keep it honest, and accept that the delete files this produces are what compaction is for. Then partition from how people will query rather than how MySQL is organised, decide that before the first load, and schedule maintenance from the first week.
Airbyte's connector catalog includes 600+ pre-built connectors, so application history can live in open formats rather than expensive storage. For the same source into a column store, see MySQL to ClickHouse, and for a document database into the same destination, MongoDB to Amazon S3 with AWS Glue.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
