Load data into Databricks from 600+ sources

Home

/

Destinations

/

Databricks

Airbyte stages Avro files inside Unity Catalog Volumes, not an external bucket, so your data stays under your catalog’s governance from the first byte to the final table. Requires a Databricks workspace with Unity Catalog enabled.

How do you load data into Databricks with Airbyte?

Four values from your workspace, one service principal, and an accepted driver license.

01

Gather your workspace details.

From your SQL warehouse or all-purpose compute cluster, collect the server hostname, HTTP path, port, and the Unity Catalog name. The catalog is the top-level catalog, not a schema or a table.
02

Create a service principal.

OAuth2 is the recommended authentication method. Create a service principal, generate a client ID and secret, and grant it access to your target catalog and schema. A personal access token also works.
03

Configure the destination.

Add Databricks as a destination in Airbyte, supply the connection details, and accept the Databricks JDBC driver terms. Then select your streams, set each to Incremental Sync - Append + Deduped where the source exposes a primary key, and run the first sync.

What happens to your grants when Airbyte syncs?

Not a philosophy. A set of specific, verifiable guarantees about the infrastructure your data runs on.

  • Staged inside the catalog.

    The connector uses Unity Catalog Volumes to stage Avro files before loading them into tables.

  • Cleaned up after.

    Purge Staging Files and Tables removes staged Avro from Volumes after loading, and is on by default.

  • Nothing retained by Airbyte.

    Airbyte does not retain customer data on its servers, so no part of the path holds a copy you do not control.

Airbyte versus Databricks Lakeflow Connect

Databricks ships its own managed ingestion, and for a subset of sources it is a reasonable choice. The question worth answering is which subset.

MethodBest forLimitsUse Airbyte instead when
Lakeflow ConnectSources on Databricks’ managed connector listBounded catalog, and it ties your ingestion layer to one lakehouseYour source is not covered, or you load into more than one destination
Auto Loader (cloudFiles)Incrementally ingesting files as they land in object storageYou own extraction and the landing processYour data starts in an API or an operational database
COPY INTOIdempotent batch loads of files already in storageSame assumption. Files firstGetting the files is the actual work
Delta Live TablesDeclarative transformation pipelinesA transformation framework, not a source connectorYou need the extraction layer beneath it
Partner ConnectDiscovering and provisioning third-party toolsA directory, not an ingestion mechanismNot applicable, this is how you would find Airbyte

Where Airbyte wins

Breadth, and not being locked to one lakehouse. 600+ sources, an open-source connector framework you can extend, and the same tooling whether you load into Databricks, Snowflake, BigQuery, or all three.

Where it does not

If Databricks is your only destination and every source you need is on the Lakeflow Connect list, staying inside one vendor is simpler and there is no strong reason to add a second.

Compared to other managed connector platforms

AirbyteFivetranDelta Live Tables
Source connectors600+Large catalogBounded managed list
Open sourceYes, and self-hostableNoNo
Runs in your own VPC or on-premYes, via Airbyte FlexLimitedNot applicable
Runs in your own VPC or on-premNoLimitedNo, Databricks only
Build your own connectorConnector Builder and CDKLimitedNo
Raman Singh headshot

Raman Singh

Tech Lead at Symend

"With our legacy framework, if one of the pipelines fails for one client, it will stop everything for the rest of our clients. But with Airbyte, things are run in parallel because of the platform’s distributed nature, which means that we can process multiple clients at the same time without impacting performance."

75%
reduction in sync times
$ 900K
in annual savings
Learn more
Sean Carver, smiling man with beard wearing blue jacket in outdoor setting with warm lighting

Sean Carver

Director of Data at PetDesk

"The real ROI is in our ability to iterate quickly, especially at our increasing scale. At the end of the day, you want a tool like that to just work. We can forget about it and know that it's configured and it's connecting and it's working. That hands-free capability is a big appeal for the platform.

20+
data sources integrated and growing
+1
FTE engineer in productivity efficiency
85%+
reduction in  data source integration time
Learn more
Man with glasses and goatee smiling at camera in professional headshot

Mondor La Grange

Head of BI and Data Engineering

"Unlike Fivetran's credit-based system that created budget uncertainty, Airbyte's pricing model allows Kuda to forecast expenses accurately and avoid surprise bills."

20+
data sources integrated and growing
+1
FTE engineer in productivity efficiency
85%+
reduction in  data source integration time
Learn more
Smiling woman with glasses and dark hair wearing patterned sweater in shopping mall

Amy Zhao

Senior Manager of Data Engineering

"What's different from Stitch Data or Informatica is the way that we can configure Airbyte connections and Airbyte entities through code. That's a huge plus to us as data engineers, because we are used to checking code and being able to manage changes from Github."

3 to 1
reduction in data integrations solutions for reduced TCO
1
week Shopify and Stripe integration with Airbyte
Learn more
Franziska Ibscher, woman with shoulder-length brown hair and white top, smiling at camera

Franziska Ibscher

Product Manager at Drivepoint

"Airbyte allows us to stay flexible while scaling from hundred-million to billion-dollar enterprise clients."

75%
of customers increased profitability
6.7%
EBITDA increase for customers
Learn more

How does your data land in Databricks?

Each stream is written directly to a final table in your configured schema. Every table carries four metadata columns: _airbyte_raw_id, _airbyte_extracted_at, _airbyte_meta, and _airbyte_generation_id.

Airbyte typeDatabricks type
stringSTRING
integerLONG
numberDECIMAL(38, 10)
booleanBOOLEAN
dateDATE
timestamp_with_timezoneTIMESTAMP, microsecond precision
timestamp_without_timezoneTIMESTAMP_NTZ, microsecond precision
time_with_timezoneSTRING, no native equivalent
objectSTRING, serialized as JSON
arraySTRING, serialized as JSON

Naming conventions

  • Schema and table names.

    Lowercased automatically. Databricks treats them as case-insensitive identifiers.

  • Column names.

    Casing from your source data is preserved.

  • Special characters.

    Escaped automatically by the connector.

The numbers

Teams running Airbyte in production

  • 0M

    pipelines synced daily

  • 6960+

    companies

  • 199%

    ROI, certified by Forrester

Raman Singh headshot

Raman Singh

Tech Lead at Symend

"With our legacy framework, if one of the pipelines fails for one client, it will stop everything for the rest of our clients. But with Airbyte, things are run in parallel because of the platform’s distributed nature, which means that we can process multiple clients at the same time without impacting performance."

75%
reduction in sync times
$ 900K
in annual savings
Learn more
Sean Carver, smiling man with beard wearing blue jacket in outdoor setting with warm lighting

Sean Carver

Director of Data at PetDesk

"The real ROI is in our ability to iterate quickly, especially at our increasing scale. At the end of the day, you want a tool like that to just work. We can forget about it and know that it's configured and it's connecting and it's working. That hands-free capability is a big appeal for the platform.

20+
data sources integrated and growing
+1
FTE engineer in productivity efficiency
85%+
reduction in  data source integration time
Learn more
Man with glasses and goatee smiling at camera in professional headshot

Mondor La Grange

Head of BI and Data Engineering

"Unlike Fivetran's credit-based system that created budget uncertainty, Airbyte's pricing model allows Kuda to forecast expenses accurately and avoid surprise bills."

20+
data sources integrated and growing
+1
FTE engineer in productivity efficiency
85%+
reduction in  data source integration time
Learn more
Smiling woman with glasses and dark hair wearing patterned sweater in shopping mall

Amy Zhao

Senior Manager of Data Engineering

"What's different from Stitch Data or Informatica is the way that we can configure Airbyte connections and Airbyte entities through code. That's a huge plus to us as data engineers, because we are used to checking code and being able to manage changes from Github."

3 to 1
reduction in data integrations solutions for reduced TCO
1
week Shopify and Stripe integration with Airbyte
Learn more
Franziska Ibscher, woman with shoulder-length brown hair and white top, smiling at camera

Franziska Ibscher

Product Manager at Drivepoint

"Airbyte allows us to stay flexible while scaling from hundred-million to billion-dollar enterprise clients."

75%
of customers increased profitability
6.7%
EBITDA increase for customers
Learn more

Where Airbyte runs

Your warehouse credentials and your data do not have to leave your environment. Airbyte does not retain customer data on its servers.

  • Airbyte Cloud

    Fully managed hosting on the Standard and Pro plans. Fastest path to a running pipeline.

  • Airbyte Flex

    Hybrid deployment, with Airbyte’s data plane running in your cloud, VPC, or on-prem environment. You get 600+ connectors and a fully managed experience, while your data never touches our servers.

  • Self-Managed

    Open source, your infrastructure, your rules.

compliance

Enterprise ready

Compliant with standards

SOC 2 Type II certified, GDPR and HIPAA support, with tools to help you meet internal and external regulatory requirements.

Uptime & SLA Guarantees

24/7 support and 99.9% availability backed by contractual SLAs and priority response times for mission-critical workloads.

What does it cost?

On the Databricks side, you pay for the SQL warehouse or all-purpose compute that runs the load, plus storage. Because staging happens in Unity Catalog Volumes, there is no separate object storage bill for the intermediate files, and Airbyte purges them after loading by default.

Core

Always free. Self-managed

For teams comfortable running open source entirely on their own.

  • Open source
  • All 600+ connectors
  • Your own infrastructure
Get started

Airbyte Cloud

Standard

Starting at $10 per month.

For practitioners who want fully managed software and prefer to be billed on data volume.

  • Fully managed cloud hosting
  • Volume-based pricing
  • Deploy quickly
Try it now

Airbyte Cloud

Pro

Capacity-based pricing.

For organizations that need scalability, governance, and security while simplifying pipeline management.

  • Fully managed cloud hosting
  • Capacity-based pricing, billed on Data Workers rather than data volume
  • Multiple workspaces
  • SSO and RBAC
  • 15-minute syncs and custom mappings
  • Premium support
Talk to sales

Enterprise Flex

Capacity-based pricing.

For enterprises in regulated industries that need full control of their data with the convenience of managed SaaS.

  • Sovereign data movement inside your own boundary
  • Available on-premises and multi-region
  • Hybrid deployment with an Airbyte-managed control plane
  • All the features of Airbyte Pro, including premium support
Talk to sales

Not sure which plan fits? Capacity-based pricing means you pay for compute capacity, not data moved, so your bill does not spike when your data does.

Frequently asked questions

Didn’t find your answer?
Please don’t hesitate to reach out.

Talk to us

If files already land in object storage, Auto Loader and COPY INTO are the closer-to-the-metal loaders. Lakeflow Connect is reasonable when every source you need is on Databricks’ managed list and Databricks is your only destination. If you need 600+ sources, or the same tooling into Databricks, Snowflake, and BigQuery, Airbyte is the extraction layer beneath those loaders.

Lakeflow Connect has a bounded catalog and ties your ingestion layer to one lakehouse. Airbyte is 600+ sources, an open-source connector framework you can extend, and the same tooling whether you load into Databricks, Snowflake, BigQuery, or all three.

No. Airbyte stages Avro files inside Unity Catalog Volumes, not an external bucket. There is no separate object storage bill for the intermediate files, and Airbyte purges them after loading by default.

Yes. This destination requires a Databricks workspace with Unity Catalog enabled. The catalog you supply is the top-level catalog, not a schema or a table.

OAuth2 is the recommended method. Create a service principal, generate a client ID and secret, and grant it access to your target catalog and schema. A personal access token also works.

Objects and arrays are serialized to STRING as JSON rather than landing in a native nested type, so plan on parsing them at query time. Timestamps carry microsecond precision, with TIMESTAMP_NTZ used where the source has no timezone.

Yes. Cloud, Flex, Self-Managed, and PyAirbyte are all supported. Flex runs the data plane in your cloud, VPC, or on-prem environment. Self-Managed runs entirely on your own infrastructure.

Other destinations

Snowflake

Airbyte clones tables instead of swapping them, so the grants you configured survive every sync. Removed source columns are kept, and oversized values are logged rather than fatal.

Load data into Snowflake

BigQuery

Airbyte stages files in Google Cloud Storage, then loads them into BigQuery with native load jobs. Warehouse compute stays on Google Cloud, and Airbyte never sits in the query path.

Load data into BigQuery

ClickHouse

Airbyte writes to ClickHouse over its native binary protocol with batch inserts. The setup needs a hostname, a port, and a user. No object storage, no staging permissions, no intermediate hop.

Load data into ClickHouse

Build with Airbyte

Ship agents and pipelines in minutes, not days.