Gong to Databricks: How to Move Your Data

Move Gong into Databricks with Airbyte. Why transcripts need a second pass against a daily cap, and how to handle verbatim conversation records.

Summarize with AI:

Moving Gong into Databricks is worth doing when you intend to process what was said. Call metadata reports anywhere; the transcripts are where the material sits, and reading several thousand sales conversations is not a job anybody is going to do by hand.

This guide covers the managed path with Airbyte. Two things shape the build: transcripts are fetched one call at a time against a daily cap, and what arrives is a verbatim record of conversations with named people in them.

Gong to Databricks at a glance:

CapabilitySupportedWhat it means for this pipeline
TranscriptsA substream of callsFetched per call, so a backfill is two passes
Daily limit10,000 callsAlongside three requests a second
Incremental behaviourOnly new callsSo steady state is cheap once the backfill lands
Private callsExcludedFrom version 1.1.0, along with their transcripts
CredentialsAdministrator onlyAnd the access key secret is shown once

Why move data from Gong to Databricks?

One situation genuinely suits this, and it is worth confirming first.

The good case is language work over transcripts. Finding which objections precede lost deals, classifying what customers actually asked for, or embedding conversations so similar ones cluster are all notebook tasks that want the text beside the compute.

The weaker case is reporting on call activity, which is counts and durations that a warehouse handles with less machinery. If governed access to conversation data is the priority, Gong to Snowflake offers controls a lakehouse does not match.

What do you need before you start?

Four things, and the first needs somebody with administrator rights:

An access key and secret from a Gong administrator. Only administrators can create them, the secret is displayed once, and the transcript read scope has to be among the permissions granted. The Gong source documentation lists the scopes.

A count of your recorded calls. It determines how long the first load takes, because transcripts are fetched per call against a daily ceiling rather than in bulk.

Permission to create Volumes in Unity Catalog. Staging goes through Avro files written into a Volume, which is separate from creating tables and worth requesting early.

Agreement from somebody accountable for recorded conversations. Transcripts contain customers and colleagues speaking, which is a different category of data from a call duration.

If your workspace restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build a Gong to Databricks pipeline in Airbyte?

Step 1: Work out how long the transcript backfill takes

Take your number of recorded calls and set it against a daily ceiling of ten thousand requests, remembering that each transcript is its own request. An organisation with two years of recordings is doing arithmetic that produces days rather than hours, and knowing that before you start turns an alarming first week into an expected one.

Step 2: Configure the Gong source

Click Sources in the left navigation, then New Source, and select Gong, following adding a source. Supply the access key, secret and a start date. Raising the concurrent stream count helps when streams are waiting on Gong to respond and does not lift the rate limit, which the connector paces itself against regardless.

Step 3: Configure the Databricks destination

Click Destinations, then New Destination, and select Databricks, following adding a destination. Supply the workspace details, warehouse or cluster, catalogue and schema. Give transcripts their own schema, since the access you want on them is narrower than on call metadata.

Step 4: Create the connection and let the first load run

Click Connections, then New connection, select your streams and a sync mode. Use incremental, since subsequent syncs fetch transcripts only for new calls and that is what makes steady state affordable. Resist restarting a slow backfill, which discards progress without lifting any limit.

Note that calls marked private are excluded along with their transcripts, so your dataset is recorded conversations minus whatever people chose to keep out of it.

Why do transcripts need a second pass?

Because the transcript stream hangs off the calls stream rather than standing alone. The connector first retrieves call metadata, then goes back and requests a transcript for each call it found, one at a time, which is how Gong's API exposes them and not something the connector chose.

That shape meets a daily ceiling of ten thousand requests and produces the defining characteristic of this pipeline. Metadata for years of calls arrives quickly; the transcripts behind them arrive over days, because each one costs a request and the allowance resets rather than stretches.

The good news is that this is a first-load problem rather than a permanent one. Once the backfill completes, incremental syncs fetch transcripts only for new calls, which for most organisations is a few dozen a day and comfortably inside any limit. Plan the beginning carefully and the steady state looks after itself.

What are you actually storing?

Verbatim records of people talking, which deserves more thought than a call duration does. A transcript is your colleagues and your customers speaking, attributed by speaker, and the customers did not choose your data platform or think about where their words would end up.

That makes two practical decisions worth making early. Keep transcripts in their own schema with access granted deliberately rather than inherited, since most analysis of call activity needs durations and participants rather than words. And check whether your recording notices and retention commitments say anything about copies, because this is one.

The lakehouse suits the work regardless. Keep the verbatim text in bronze so a model's output can always be checked against what was said, and put classifications, extracted themes and embeddings in silver beside it. A theme label nobody can trace back to a sentence is not evidence, and on this data being able to show your working matters more than usual.

Frequently asked questions

Why is the first sync taking days?

Transcripts are fetched one call at a time against a daily limit of ten thousand requests, so a large history spans several days. Later syncs cover only new calls.

Will more threads speed it up?

Not against the rate limit, which the connector paces itself to regardless. More threads help only when streams are waiting on Gong to respond.

Are all calls included?

No. From version 1.1.0 calls marked private are excluded, and because transcripts hang off calls, their transcripts are excluded too.

Who can create the credentials?

A Gong administrator, and the access key secret is shown once, so capture it at creation rather than expecting to retrieve it.

Can I do this without writing code?

The pipeline, yes. The models that read transcripts are the reason to choose this destination, and they are code by definition.

Get your Gong data into Databricks

Do the arithmetic on your call count against a daily limit before starting, because transcripts arrive one request at a time and the first load is measured in days. Get credentials from an administrator and capture the secret immediately. Then keep transcripts in their own schema with deliberate access, and retain the verbatim text in bronze so anything a model concludes can be traced to a sentence.

Airbyte's connector catalog includes 600+ pre-built connectors, so sales conversations can be analysed at a scale nobody could read. For the same source into a warehouse, see Gong to BigQuery, and for survey responses into the same destination, Surveymonkey to Databricks.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.