Connectors

/

Databricks Lakehouse

Data replication

Connect Databricks Lakehouse with any data sources.

Databricks is an American enterprise software company founded by the creators of Apache Spark. Databricks combines data warehouses and data lakes into a lakehouse architecture.

  • Certified
  • Alpha
Databricks Lakehouse

Everything Databricks Lakehouse can do in Airbyte

  • Load from any source

    Move data into Databricks Lakehouse from 600+ Airbyte sources on a schedule you control.

  • Incremental syncs

    Pull only the records that changed since the last run instead of reloading everything.

  • Cloud or self-hosted

    Run the Databricks Lakehouse connector on Cloud, Self-Managed Enterprise.

What to know before you load into Databricks Lakehouse

  • Support levelCertified
  • Available onCloud, Self-Managed Enterprise
  • Connector version4.0.3
  • Release stageAlpha

Sync capabilities

  • Full Refresh SyncSupported
  • Incremental SyncSupported
  • Sources600+ Airbyte connectors

Set up in 11 steps

  1. First, navigate to the Airbyte website and log in to your account.
  2. Once you are logged in, click on the "Destinations" tab on the left-hand side of the screen.
  3. Scroll down until you find the "Databricks Lakehouse" connector and click on it.
  4. You will be prompted to enter your Databricks Lakehouse credentials, including your account name, personal access token, and workspace ID.
  5. Once you have entered your credentials, click on the "Test" button to ensure that the connection is successful.
  6. If the test is successful, click on the "Save" button to save your Databricks Lakehouse destination connector settings.
  7. You can now use the Databricks Lakehouse connector to transfer data from your source connectors to your Databricks Lakehouse destination.
  8. To set up a data transfer, navigate to the "Sources" tab and select the source connector that you want to use.
  9. Follow the prompts to enter your source connector credentials and configure your data transfer settings.
  10. Once you have configured your source connector, select the Databricks Lakehouse connector as your destination and follow the prompts to configure your data transfer settings.
  11. Click on the "Run" button to initiate the data transfer.

Common questions

Didn't find your answer?
Please don't hesitate to reach out.

Talk to sales

ETL, an acronym for Extract, Transform, Load, is a vital data integration process. It involves extracting data from diverse sources, transforming it into a usable format, and loading it into a database, data warehouse or data lake. This process enables meaningful data analysis, enhancing business intelligence.

Databricks Lakehouse provides access to a wide range of data types, including: Structured data (organized into tables with defined columns and data types, such as CSV, JSON, and Avro files); Semi-structured data (some structure, but not necessarily a fixed schema, such as XML and JSON files); Unstructured data (no predefined structure, such as text, images, and videos); Time-series data (organized by time, such as stock prices, weather data, and sensor readings); Geospatial data (related to geographic locations, such as maps, GPS coordinates, and spatial databases); Machine learning data (used to train machine learning models, such as labeled datasets and feature vectors); and Streaming data (generated in real-time, such as social media feeds, IoT sensor data, and log files). Overall, Databricks Lakehouse's API provides access to a wide range of data types, making it a powerful tool for data analysis and machine learning.

1. First, navigate to the Airbyte website and log in to your account.
2. Once you are logged in, click on the "Destinations" tab on the left-hand side of the screen.
3. Scroll down until you find the "Databricks Lakehouse" connector and click on it.
4. You will be prompted to enter your Databricks Lakehouse credentials, including your account name, personal access token, and workspace ID.
5. Once you have entered your credentials, click on the "Test" button to ensure that the connection is successful.
6. If the test is successful, click on the "Save" button to save your Databricks Lakehouse destination connector settings.
7. You can now use the Databricks Lakehouse connector to transfer data from your source connectors to your Databricks Lakehouse destination.
8. To set up a data transfer, navigate to the "Sources" tab and select the source connector that you want to use.
9. Follow the prompts to enter your source connector credentials and configure your data transfer settings.
10. Once you have configured your source connector, select the Databricks Lakehouse connector as your destination and follow the prompts to configure your data transfer settings.
11. Click on the "Run" button to initiate the data transfer.

The most prominent ETL tools to transfer data to include: Airbyte, Fivetran, StitchData, Matillion, Talend Data Integration. These tools help in extracting data from various sources (APIs, databases, and more), transforming it efficiently, and loading it into and other databases, data warehouses and data lakes, enhancing data management capabilities.

ELT, standing for Extract, Load, Transform, is a modern take on the traditional ETL data integration process. In ELT, data is first extracted from various sources, loaded directly into a data warehouse, and then transformed. This approach enhances data processing speed, analytical flexibility and autonomy.

ETL and ELT are critical data integration strategies with key differences. ETL (Extract, Transform, Load) transforms data before loading, ideal for structured data. In contrast, ELT (Extract, Load, Transform) loads data before transformation, perfect for processing large, diverse data sets in modern data warehouses. ELT is becoming the new standard as it offers a lot more flexibility and autonomy to data analysts.

Start moving Databricks Lakehouse data today

Free for 14 days on Airbyte Cloud. Set up the Databricks Lakehouse connector once and let Airbyte keep it in sync.