Google Ads to Databricks: How to Move Your Data

Move Google Ads into Databricks with Airbyte. Why the developer token needs Google's approval, why 37-month retention is silent, and how to keep granular rows.

Summarize with AI:

Moving Google Ads into Databricks gives you granular advertising history somewhere it can be kept and modelled properly. Google Ads reports capably within its own interface and expires the detailed rows behind those reports on a schedule, which is awkward when the analysis you want reaches back four years.

This guide covers the managed path with Airbyte. Two things shape the build: you cannot start today because access needs approval, and granular data expires after roughly 37 months, which decides what this destination is really for.

Google Ads to Databricks at a glance:

CapabilitySupportedWhat it means for this pipeline
Developer tokenNeeds approvalRequires a Manager account and a request to Google
Granular retentionAbout 37 monthsOlder report rows are skipped silently rather than erroring
Conversion window14 days defaultRecent figures keep maturing, so they are provisional
Manager accessExtra fieldA login customer ID is needed when accessing via a manager
Volume ceilingNone in practiceGranular rows can be kept indefinitely rather than aggregated

Why move data from Google Ads to Databricks?

Two situations account for most of these pipelines.

The first is keeping detail rather than summaries. Most organisations retain advertising history as monthly aggregates because that is what fits comfortably in a warehouse, and then discover that the question they now want to answer needs the keyword-level or ad-level rows they threw away. A lakehouse has no practical ceiling, so keeping the granular version is affordable.

The second is modelling spend against outcomes that live elsewhere, joining campaigns to revenue, retention and support cost. If the requirement is straightforward reporting with no table maintenance to operate, a warehouse asks less of you and Google Ads to BigQuery covers that comfortably.

What do you need before you start?

Four things, and the first one has a queue attached:

A developer token, which Google has to approve. It is requested from a Manager account and reviewed rather than issued instantly, so this is a lead time rather than a form. The Google Ads source documentation covers the request and the credentials that go with it.

Your customer identifiers, plus a login customer ID if relevant. Accessing accounts through a manager requires that extra field, and omitting it produces an access failure that looks like a permissions problem rather than a missing parameter.

Permission to create Volumes in Unity Catalog. Staging goes through Avro files written into a Volume, and that permission is separate from creating tables. Request it while you are waiting for the developer token, since the two waits can overlap.

An agreed conversion window. It defaults to fourteen days, and it decides what counts as a conversion. Agree it with whoever compares channels rather than inheriting a default nobody chose.

If your workspace restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build a Google Ads to Databricks pipeline in Airbyte?

Step 1: Request the developer token before anything else

This is the step with a queue in it, so start it on day one and do everything else while you wait. The request comes from a Manager account and Google reviews it, which means a project planned as a week of work can spend most of that week waiting on somebody else. Teams routinely leave this until the configuration stage and then discover the schedule was never theirs to control.

Step 2: Configure the Google Ads source

Click Sources in the left navigation, then New Source, and select Google Ads, following adding a source. Supply the developer token, OAuth credentials, customer identifiers, start date and conversion window, adding the login customer ID if you are reaching accounts through a manager. Set the start date knowing it cannot usefully reach beyond the retention limit.

Step 3: Configure the Databricks destination

Click Destinations, then New Destination, and select Databricks, following adding a destination. Supply the workspace details, warehouse or cluster, catalogue and schema. Report rows are wide and sometimes nested, and Spark handles that without complaint, so nothing needs flattening on the way in.

Step 4: Create the connection and treat recent days as provisional

Click Connections, then New connection, select your streams and a sync mode. Conversions continue being credited to earlier clicks for the length of your conversion window, so the last fortnight of figures will keep changing. Handle that in the silver layer rather than letting a dashboard present a maturing number as final.

Daily is right. Attribution settles over weeks, so syncing more often produces repeated work rather than fresher insight.

Why can't you start this today?

Because the developer token is granted rather than generated. You request it from a Google Ads Manager account and Google reviews the request, which is a genuine wait rather than a form you fill in. No amount of preparation elsewhere shortens it, and the pipeline cannot be tested at all until it arrives.

The token's access level then governs how many operations you get per day, which is worth understanding because a basic level suits a small advertiser and constrains an agency managing many accounts. If your daily allowance turns out to be the limiting factor, that is a conversation with Google rather than a setting to adjust.

The login customer ID is the smaller trap alongside it. When you reach accounts through a manager rather than directly, that field is required, and leaving it out produces an access error indistinguishable from a permissions problem. People spend an afternoon checking account permissions that were correct all along, so try the field before you try the permissions.

Why does 37 months make a lakehouse the right home?

Because granular data expires and this is the destination that can afford to keep all of it. Google Ads retains detailed report rows for roughly 37 months, and the connector skips anything older silently rather than raising an error, so a start date reaching back five years produces a dataset quietly beginning three years ago.

That silence is the part to plan around. A backfill appears to succeed, the tables fill, and nobody notices the missing years until somebody asks a question spanning them. Check the earliest date actually present in the data after the first sync rather than assuming the start date you typed was honoured.

Beyond that window, your tables are the only copy, which is where the absence of a volume ceiling earns its keep. Most organisations aggregate advertising history to save space and lose the ability to ask keyword-level questions about last year. Here you can keep the raw rows in bronze indefinitely and build whatever aggregates the business currently wants in silver, so a new question about old data is a query rather than an apology.

Frequently asked questions

How long does the developer token take?

Long enough to matter. It is requested from a Manager account and approved by Google rather than issued automatically, so start the request before anything else and plan the rest around it.

Why does my backfill start later than I asked for?

Granular retention is about 37 months and older rows are skipped silently. Check the earliest date present after the first sync rather than trusting the start date you set.

I get an access error but the permissions look right.

Check whether you need a login customer ID, which is required when reaching accounts through a manager. Its absence produces an error that reads like a permissions failure.

Why do last week's conversions keep rising?

The conversion window, which defaults to fourteen days. Conversions continue being credited to earlier clicks, so recent figures mature rather than staying fixed.

Can I do this without writing code?

The pipeline, yes. The silver layer handling maturing conversion figures and building the aggregates your reporting uses is modelling work, and it is where this arrangement pays off.

Get your Google Ads data into Databricks

Request the developer token on day one, because Google's review is the part of the timeline you do not control. Add the login customer ID if you reach accounts through a manager, and check the earliest date in your data after the first sync, since anything beyond about 37 months is skipped without complaint. Then keep the granular rows in bronze rather than aggregating them away, because that is the advantage this destination gives you over a warehouse.

Airbyte's connector catalog includes 600+ pre-built connectors, so advertising spend can be modelled against the outcomes it produced. For the same source into a warehouse, see Google Ads to Snowflake, and for mobile attribution into the same destination, AppsFlyer to Databricks.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.