Metabase to Amazon S3 with AWS Glue: How to Move Your Data
Move Metabase into S3 with AWS Glue using Airbyte. Why a tiny catalogue in Iceberg is a standardisation choice, and why small files accumulate anyway.

Moving Metabase into Amazon S3 with AWS Glue puts your reporting catalogue into the same open format as everything else on your platform. That is the whole argument, and it is worth being clear that it is an architectural one rather than a technical necessity.
This guide covers the managed path with Airbyte. Two things shape the build: a Metabase catalogue is tiny and Iceberg is built for the opposite, and a small dataset synced regularly creates a maintenance problem that has nothing to do with its size.
Metabase to Amazon S3 with AWS Glue at a glance:
Why move data from Metabase to Amazon S3 with AWS Glue?
One situation genuinely justifies this, and the other is worth naming honestly.
The good case is consistency. If your organisation has standardised on Iceberg tables in S3, with one catalogue, one permission model and one set of engines, then putting your reporting metadata anywhere else creates an exception somebody has to maintain. Uniformity has real value and it is a legitimate reason to choose a format the data does not need.
The weaker case is treating this as a data volume decision, because a Metabase catalogue is small by any measure and Iceberg exists to make enormous tables manageable. If you are not already committed to the format, Metabase to BigQuery gets you the same answers with nothing to maintain.
What do you need before you start?
Four things, and the first is a question about your platform rather than this pipeline:
A reason to be in Iceberg at all. If the answer is that everything else lives there, proceed. If it is that Iceberg seemed modern, a warehouse will serve this dataset better for less operational effort.
Metabase credentials, and a plan for keeping them working. A session token expires after roughly a fortnight, while username and password authentication creates a session per query and can trigger security alerts. The Metabase source documentation covers both.
An S3 bucket and a Glue catalog. The pipeline writes data files to one and registers tables in the other. Keep namespace and table names alphanumeric with underscores, since Glue rewrites anything else.
A compaction schedule, even though the data is small. This is the part people skip on a tiny dataset, and it is the part that matters most here.
If your network or Metabase instance restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow lists before you begin.
How do you build a Metabase to Amazon S3 with AWS Glue pipeline in Airbyte?
Step 1: Check this is a standardisation decision
Ask why Iceberg, and accept consistency as a complete answer if that is the real one. A platform where every dataset is an Iceberg table with one catalogue and one access model is easier to run than one with a warehouse exception for reporting metadata, and that is worth the maintenance. What does not justify it is the size or shape of this data, which would be perfectly happy in a small table almost anywhere.
Step 2: Configure the Metabase source
Click Sources in the left navigation, then New Source, and select Metabase, following adding a source. Supply your instance URL and credentials. Expect the definitions of questions, dashboards and collections rather than the rows those questions return, which live in whichever database Metabase queries.
Step 3: Configure the S3 Data Lake destination
Click Destinations, then New Destination, and select the S3 Data Lake, following adding a destination. Choose AWS Glue as the catalog and supply the bucket, region and credentials. Name things in plain letters, digits and underscores so the catalog reads the way you intended.
Step 4: Create the connection and schedule it slowly
Click Connections, then New connection, select your streams and a sync mode. Daily is generous for a catalogue that changes when somebody builds a dashboard, and on this destination a gentle schedule is not only sufficient but actively helpful.
Then set up compaction, which on a dataset this small feels unnecessary and is not.
Is a lake format right for a catalogue?
On its own merits, no. Iceberg earns its complexity on tables large enough that partition pruning, snapshot isolation and schema evolution solve real problems. A Metabase catalogue is hundreds or thousands of rows describing questions and dashboards, and none of those capabilities is doing anything for it.
That is not an argument against doing it, because consistency is a genuine engineering value. A platform where one dataset lives somewhere else is a platform with an exception, and exceptions cost attention every time somebody grants access, audits storage or writes a query that needs two engines. Standardising is a reasonable reason to put small data in a big format.
What matters is being honest about which reason applies. If you are standardising, proceed and accept the maintenance. If you reached for Iceberg because it is what your platform team uses for everything else without that being a deliberate policy, a small table in a warehouse answers the same governance questions with nothing to schedule and nobody to remind.
Why does a small dataset still need compaction?
Because file count follows sync frequency rather than data volume. Every sync writes files, and a table receiving a few hundred rows daily accumulates files at exactly the same rate as one receiving millions. After a year you have a year's worth of files describing a dataset that would fit comfortably in one of them.
This is the small files problem, and it is counterintuitive precisely because the data is tiny. Nobody schedules maintenance for a table measured in megabytes, so it gets skipped, and then reads slow down for reasons that look absurd given the size. The metadata overhead of tracking thousands of files starts to dominate the cost of reading the files themselves.
Two things fix it cheaply. Schedule compaction and snapshot expiry as you would for any Iceberg table, since the job is trivial at this size and the omission is what causes trouble. And sync gently, because the frequency you choose is the thing generating files: daily on a catalogue that changes weekly produces seven times the files for no additional information.
Frequently asked questions
Does this give me the data behind my dashboards?
No. It carries the catalogue, meaning the definitions of questions, dashboards and collections. The rows those questions return stay in whichever database Metabase queries.
Is Iceberg overkill for this?
On the data's own merits, yes. As a standardisation decision on a platform where everything else is an Iceberg table, it is entirely reasonable.
Why are reads slow on such a small table?
Almost certainly accumulated small files, since file count follows sync frequency rather than volume. Schedule compaction and sync less often.
Why did the pipeline stop after two weeks?
The session token expired, since those last around a fortnight. Rotate it, or self-host and extend the session duration.
Can I do this without writing code?
The pipeline, yes. Compaction and snapshot expiry are jobs you schedule, and on this pairing they are the difference between a table that stays fast and one that quietly does not.
Get your Metabase data into Amazon S3 with AWS Glue
Be honest about why you are here, because consistency across a platform standardised on Iceberg is a good reason and the data's size is not. Plan for the session token expiring. Then schedule compaction despite the dataset being tiny, since file count follows how often you sync rather than how much arrives, and sync gently for the same reason. A catalogue that changes weekly does not need reading every day.
Airbyte's connector catalog includes 600+ pre-built connectors, so a reporting estate can be governed in whatever format your platform has settled on. For the same source into an operational database, see Metabase to PostgreSQL, and for a larger source into the same destination, Hubspot to Amazon S3 with AWS Glue.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
