Appsflyer to Amazon S3 with AWS Glue: How to Move Your Data
Move AppsFlyer into S3 with AWS Glue using Airbyte. Whether apps share a table, and why a 90 day source window makes your lake irreplaceable.

Moving AppsFlyer into Amazon S3 with AWS Glue builds an attribution archive in an open format. AppsFlyer serves raw reports for a limited window, so a lake that keeps accumulating is the only place a question about last year can be answered.
This guide covers the managed path with Airbyte. Two things shape the build: a portfolio of apps means many sources and a decision about how they land, and your tables become the only copy of data nobody can fetch again.
Appsflyer to Amazon S3 with AWS Glue at a glance:
Why move data from Appsflyer to Amazon S3 with AWS Glue?
Two situations account for most of these pipelines.
The first is retention at a sensible price. Attribution data is worth keeping for years and consulted occasionally, so storage cost matters more than query latency, and Iceberg tables on object storage hold it cheaply while staying readable by several engines.
The second is feeding transformation work across a portfolio of apps. If this data needs governed access controls because of what device-level records contain, Appsflyer to Snowflake offers masking and row access policies a lake does not match.
What do you need before you start?
Four things, and the first determines how many pipelines this project involves:
A list of your apps. The app identifier is a single field, so one source covers one app and a portfolio across two platforms doubles the count. The AppsFlyer source documentation covers the fields and the available streams.
An API token and your project timezone. Only an account admin can create the token, and AppsFlyer reissues it when that admin changes. The timezone should match the app settings in the AppsFlyer console or your daily figures shift.
Awareness of the ninety-day raw data window. A start date further back is capped rather than refused, so a long backfill can appear to succeed while quietly covering three months.
An S3 bucket, a Glue catalog and a maintenance plan. Keep namespace and table names alphanumeric with underscores, since Glue rewrites anything else, and think about maintenance differently here for reasons below.
If your network restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow list before you begin.
How do you build an Appsflyer to Amazon S3 with AWS Glue pipeline in Airbyte?
Step 1: Decide whether apps share a table
Work out now whether installs from six apps belong in one table with a column naming the app, or in six tables. One table makes portfolio questions trivial and per-app access harder; separate tables do the reverse. This is much easier to settle before anything is written than after a year of data has accumulated under whichever arrangement happened first.
Step 2: Configure the AppsFlyer source
Click Sources in the left navigation, then New Source, and select AppsFlyer, following adding a source. Supply the API token, app identifier, start date and timezone, then repeat per app. Name each source after the app and platform, because a workspace of near-identical AppsFlyer sources is easy to create and unpleasant to maintain.
Step 3: Configure the S3 Data Lake destination
Click Destinations, then New Destination, and select the S3 Data Lake, following adding a destination. Choose AWS Glue as the catalog and supply the bucket, region and credentials. If several sources are landing in one table, they all need the same namespace and table naming.
Step 4: Create the connections and reconcile a day
Click Connections, then New connection for each source, selecting your streams and a sync mode. Then compare one settled day against the AppsFlyer dashboard, because a mismatch is almost always the timezone rather than a broken pipeline, and it is far easier to diagnose early.
Alert on failure from the first week, since a pipeline stopped for three months has lost data permanently rather than fallen behind.
One table or one per app?
A question this source forces on you, because the app identifier is a single field and every app needs its own source. Six apps across two platforms is a dozen configurations feeding a lake, and where they land is a design decision rather than something the pipeline settles.
One table per stream with a column naming the app is usually the better answer. Portfolio questions become ordinary filters rather than unions across a dozen tables, partitioning by date and app gives queries something to prune on, and maintenance is a handful of jobs rather than a dozen. It also keeps the table count manageable as the portfolio grows.
Separate tables earn their place when access needs to differ by app, since granting a team one table is simpler than filtering a shared one, or when different apps genuinely carry different fields. Decide on that basis rather than by default, and remember that whichever you choose becomes progressively harder to change as the archive grows.
What changes when your table is the only copy?
Your appetite for risk in maintenance, which is the genuinely different thing about this pairing. Raw reports reach back about ninety days, so anything older exists solely in what you collected, and the usual comfort of a lake, that you can always resync from the source, does not apply beyond that window.
Compaction and snapshot expiry remain necessary, and they operate on files no longer referenced rather than on your current data, so routine maintenance is safe. What is not safe is anything that rewrites or removes rows: a mistaken delete, an overzealous backfill or a table dropped during a reorganisation cannot be repaired by running the pipeline again.
So treat these tables with more care than a lake usually warrants. Keep a retention window on snapshots long enough that a mistake is noticed before the ability to roll back expires, restrict who can drop or rewrite them, and alert on failure so nobody discovers a three-month gap in a quarterly review. The pipeline is cheap; the data it collects is not replaceable.
Frequently asked questions
Can I backfill more than ninety days?
No. Raw reports reach back about ninety days and an older start date is capped silently, so do not assume a long backfill succeeded because nothing complained.
My totals disagree with the AppsFlyer dashboard.
Check the timezone setting against your app settings first, since a mismatch moves events between days and is the usual explanation.
Should each app have its own table?
Usually not. One table per stream with an app column makes portfolio questions easier, unless access needs to differ by app or the apps carry genuinely different fields.
The pipeline stopped working after a staff change.
The API token is tied to the account admin, and AppsFlyer reissues it when that admin changes. Update the configuration with the new token.
Can I do this without writing code?
The pipelines, yes, though there is one per app. Compaction and snapshot expiry are jobs you schedule, and here they deserve more thought than usual.
Get your Appsflyer data into Amazon S3 with AWS Glue
Count your apps, since each needs its own source, and decide whether they share a table before anything is written. Match the timezone to your app settings and reconcile a settled day against the dashboard. Then treat these tables as irreplaceable, because raw reports reach back only ninety days: keep a generous snapshot retention, restrict who can rewrite them, and alert on failure.
Airbyte's connector catalog includes 600+ pre-built connectors, so attribution history can outlive the window the platform keeps. For the same source into a lakehouse, see Appsflyer to Databricks, and for a CRM into the same destination, Hubspot to Amazon S3 with AWS Glue.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
