Amplitude to Snowflake: How to Move Your Data
Move Amplitude into Snowflake with Airbyte. Why cost-based and size-based limits fail differently, and why computed metrics belong apart from raw events.

Moving Amplitude into Snowflake puts product behaviour beside the commercial data that gives it meaning. Amplitude knows what users did and nothing about what any of it was worth, so questions about which behaviours predict renewal need events and contracts together.
This guide covers the managed path with Airbyte. Two things shape the build: the connector spans several Amplitude APIs whose limits work in completely different ways, and some of the streams are Amplitude's own calculations rather than raw data.
Amplitude to Snowflake at a glance:
Why move data from Amplitude to Snowflake?
Two situations account for most of these pipelines.
The first is joining behaviour to value, which Amplitude cannot do because it has never seen a contract or a support ticket. Whether a feature predicts retention, and what that retention is worth, is a question for a warehouse holding both sides.
The second is governed access, since event data identifies individual users and Snowflake's controls suit that. If your priority is interactive speed across very large event volumes rather than joins and governance, Amplitude to ClickHouse is the better shape.
What do you need before you start?
Four things, and the second is the one that produces a confusing failure:
An API key and secret key. Both found in your Amplitude project settings, and both are per project rather than per account, so several projects means several sources. The Amplitude source documentation lists every field.
The correct data region. The setting defaults to the standard server, so a project hosted in the EU data centre needs the EU residency option selected, and getting it wrong finds no data rather than explaining itself.
A rough figure for daily event volume. It determines your request time range, because exports have a hard size ceiling and a busy product passes it comfortably within twenty-four hours.
A view on which streams you actually want. Some carry raw events and others carry Amplitude's own computed metrics, and mixing them without noticing is the most common way this dataset misleads people.
If your Snowflake account restricts inbound traffic by IP, add the Airbyte Cloud IP addresses to the network policy before you begin.
How do you build an Amplitude to Snowflake pipeline in Airbyte?
Step 1: Decide which of the four APIs you need
This connector spans event exports, dashboard metrics, chart annotations and behavioural cohorts, and those are four quite different kinds of thing. Work out which your analysis needs before selecting streams, because the answer determines both how long your syncs take and whether you end up with two versions of the same measure in one schema. Most projects want the events and rather less of the rest than they first assume.
Step 2: Configure the Amplitude source
Click Sources in the left navigation, then New Source, and select Amplitude, following adding a source. Supply the API key, secret key, data region and start date, then set the request time range from your volume estimate. If the connection check itself fails, try disabling the option that groups active users by country, which the documentation flags for exactly that symptom.
Step 3: Configure the Snowflake destination
Click Destinations, then New Destination, and select Snowflake, following adding a destination. Supply the account identifier, warehouse, database, schema and role. Consider separate schemas for raw events and computed metrics, since keeping them visually apart is the cheapest guard against the confusion described below.
Step 4: Create the connection and watch the export stream
Click Connections, then New connection, select your streams and a sync mode. If the events stream errors or times out, reduce the request time range so each request covers fewer hours. That is the single adjustment this pipeline usually needs, and it is the one nobody thinks of first.
Daily is right for product analytics, and the dashboard streams will finish long before the events do.
Why do the two halves fail so differently?
Because Amplitude meters them on unrelated principles. The dashboard API applies cost-based limits: a budget spent per hour, a smaller budget per short burst window, and a ceiling on concurrent requests. The connector tracks what each request costs and throttles itself to stay inside that, which means it slows down and keeps working.
The export API used by the events stream has a size ceiling instead, capping each export at four gigabytes and returning an error when a request exceeds it. That is a failure rather than a delay, and no amount of throttling avoids it, because the problem is how much data one request asked for rather than how often you asked.
So the two halves need different responses from you, which is exactly none and one respectively. The dashboard streams look after themselves. The events stream needs a request time range narrow enough that a single window stays under the cap, and on a high-traffic product that means considerably less than the twenty-four hour default. Timeouts on large volumes have the same remedy.
Which streams are raw and which are Amplitude's opinion?
A distinction worth drawing clearly, because both arrive as tables and look equally authoritative. Event exports are raw records of what happened. Active user counts and average session length are metrics Amplitude computed, using Amplitude's definitions of what an active user is and where a session begins and ends.
The trouble starts when somebody computes their own active user count from the events and compares it to the one Amplitude provided. The numbers will differ, because the definitions differ, and there is no bug to find. An afternoon disappears into reconciling two figures that were never meant to agree.
So decide which is authoritative for each question and write that down beside the tables. Amplitude's computed metrics are the right answer when you want to match what the product team sees in Amplitude; your own derivation from raw events is right when you need a definition that matches how the rest of the business counts things. Keeping the two in separate schemas makes the distinction visible before anybody joins them by accident.
Frequently asked questions
The events stream errors or times out.
Reduce the request time range so each export covers fewer hours. Exports are capped at four gigabytes and large windows also time out, and this is the fix for both.
Do I need to manage rate limits myself?
For the dashboard API, no. It uses cost-based budgets and the connector tracks per-request costs and throttles automatically. The export size ceiling is the part you control.
The connection check fails.
Check the data region first, then try disabling the option that groups active users by country, which the documentation names as a cause of check and fetch failures.
My active user count disagrees with Amplitude's.
That is expected. Amplitude's figure uses its own definitions, and yours derived from raw events uses yours. Decide which is authoritative per question rather than reconciling them.
Can I do this without writing code?
The pipeline, yes. The views joining events to commercial data, and the documentation of which metric is authoritative, are the work that makes this usable.
Get your Amplitude data into Snowflake
Choose which of the four APIs you genuinely need, since that decides both your sync time and how confusing the result is. Set the data region deliberately. Then size the request time range from your daily event volume, because the export cap fails rather than throttles, and keep Amplitude's computed metrics in a separate schema from your raw events so nobody spends an afternoon reconciling two numbers that were never meant to match.
Airbyte's connector catalog includes 600+ pre-built connectors, so product behaviour can be measured against what it was worth. For the same source into a warehouse on another platform, see Amplitude to BigQuery, and for another product analytics platform into the same destination, Pendo to Snowflake.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
