Salesforce to Databricks: How to Move Your Data

Move Salesforce into Databricks with Airbyte. Why a successful sync can be incomplete, and why formula field values go stale without any error.

Summarize with AI:

Moving Salesforce into Databricks gives commercial data somewhere it can be modelled properly and joined to everything else. A CRM records what sales people entered; understanding what it means alongside product usage, support load and delivery cost is work for a lakehouse.

This guide covers the managed path with Airbyte. Two things shape the build: a sync can report success while stopping early, and formula fields go stale in a way nothing in the pipeline will tell you about.

Salesforce to Databricks at a glance:

CapabilitySupportedWhat it means for this pipeline
Daily rate limitsEnd the sync earlyWith a success status, resuming on the next run
ResumptionIncremental onlyWhich is why deduped history is the recommended mode
Formula fieldsOutputs onlyA changed formula needs a stream reset and backfill
API selectionAutomaticBulk by default, REST where Bulk cannot cope
DeletionsFlagged, not removedDeleted records arrive marked rather than disappearing

Why move data from Salesforce to Databricks?

Two situations account for most of these pipelines.

The first is modelling a commercial picture that spans systems. Pipeline conversion, customer health and cost to serve all need CRM records beside product events and support history, and deriving those measures is transformation work that suits a lakehouse.

The second is keeping the full history of how records changed, which a CRM overwrites. If your need is governed reporting with fine-grained access controls rather than heavy transformation, Salesforce to Snowflake suits that better and asks less of you operationally.

What do you need before you start?

Four things, and the second is a decision the connector's behaviour depends on:

A Salesforce user whose permissions match what you want. Streams are discovered dynamically from what that user can read and what is queryable, so a missing object is usually a permissions question. The Salesforce source documentation covers authentication and the object filters.

A commitment to incremental deduped history. Recommended for a reason covered below, and not simply a preference: the connector's behaviour at the daily rate limit depends on it.

Permission to create Volumes in Unity Catalog. Staging goes through Avro files written into a Volume, which is separate from creating tables and worth requesting early.

A list of the objects you actually need. Salesforce exposes a great many standard and custom objects, each as its own stream, and taking everything makes the daily limit arrive sooner.

If your workspace restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build a Salesforce to Databricks pipeline in Airbyte?

Step 1: Choose incremental deduped history from the start

This is not the usual advice about efficiency. Salesforce enforces daily rate limits, and when the connector reaches one it ends the sync early, reports success and picks up where it stopped on the next run. That resumption only works for incremental syncs, so the mode you choose determines whether an interrupted sync recovers gracefully or starts again tomorrow and stops at the same place.

Step 2: Configure the Salesforce source

Click Sources in the left navigation, then New Source, and select Salesforce, following adding a source. Authenticate, set a start date, and use the object filters to narrow what you sync. A blank start date replicates the last two years, which is a reasonable default and worth choosing rather than inheriting.

Step 3: Configure the Databricks destination

Click Destinations, then New Destination, and select Databricks, following adding a destination. Supply the workspace details, warehouse or cluster, catalogue and schema. Keep the landing tables faithful to what Salesforce returned, since the modelling belongs in a silver layer rather than on the way in.

Step 4: Create the connection and check for early finishes

Click Connections, then New connection, select your streams and a sync mode. For the first week, compare record counts against Salesforce rather than trusting a green run, because a successful status here does not always mean everything arrived.

Note that deleted records arrive flagged rather than removed, so your silver layer needs to decide what to do with them.

Why might a successful sync be incomplete?

Because Salesforce meters by day and the connector handles that by stopping rather than failing. When the daily allowance is exhausted it ends the sync early, marks it successful and resumes from that point next time. That is a sensible design, since a partial load you can continue beats a failure you must repeat, and it means a green run is not proof of a complete one.

The resumption is the part with a condition attached. Picking up where it left off works for incremental syncs, so a full refresh that hits the limit simply stops and starts again from the beginning next time, potentially never finishing on a large org. That is why deduped history is the recommended mode rather than merely the tidiest.

Selecting fewer objects helps more than anything else, because each one is its own stream and they all draw on the same daily budget. The connector chooses between the bulk and REST interfaces per object automatically, preferring bulk and falling back where an object or a column type is unsupported, and REST-served objects cost more of your quota. So a narrower selection is both faster and less likely to stop early.

What happens when a formula changes?

Your warehouse keeps the old answers, quietly. Salesforce formula fields are computed, and the connector syncs their output rather than the formula itself. If somebody edits a formula and nothing else on a record changes, that record has no reason to appear in an incremental sync, so the stored value stays as it was.

This is a genuinely nasty failure mode because it is invisible from both ends. Salesforce shows the new values, your tables show the old ones, and nothing errors. Formula fields are often exactly the ones business users care about, carrying scores, tiers and derived amounts that feed reporting, so the divergence tends to be found by somebody comparing a dashboard to the CRM.

The remedy is a stream reset and a historical backfill, which is where this destination helps. A lakehouse holds full history cheaply and re-running a backfill is routine rather than an event, so the cost of correcting is mostly the API budget. What matters more is hearing about it, so ask your Salesforce administrators to tell you when a formula changes, and treat that message as a trigger rather than a note.

Frequently asked questions

My sync succeeded but records are missing.

It probably hit the daily rate limit and ended early with a success status. Incremental syncs resume from that point on the next run.

Why is deduped history recommended?

Because resumption after an early finish only works for incremental syncs. A full refresh that hits the limit restarts from the beginning next time.

A formula field value looks out of date.

If the formula changed and no other field on the record did, nothing prompted a resync. Reset the stream and run a historical backfill.

Do deleted records disappear?

No. For streams that support it, deleted records arrive flagged as deleted, so your models need to filter them rather than assuming absence means removal.

Can I do this without writing code?

The pipeline, yes. The silver layer handling deleted flags and the checks comparing counts against Salesforce are work worth doing deliberately.

Get your Salesforce data into Databricks

Choose incremental deduped history from the start, because the connector's behaviour at the daily rate limit depends on it and a full refresh may never finish on a large org. Select objects narrowly for the same reason. Then arrange to hear when Salesforce formulas change, since their outputs go stale with no error anywhere, and reset the affected stream when they do.

Airbyte's connector catalog includes 600+ pre-built connectors, so commercial data can be modelled beside everything that explains it. For the same source into a warehouse, see Salesforce to Snowflake, and for workspace content into the same destination, Notion to Databricks.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.