GCS Data Lake
Connect GCS Data Lake with any data sources.
- Certified
- Generally available
- 600+Sources
Everything GCS Data Lake can do in Airbyte
Load from any source
Move data into GCS Data Lake from 600+ Airbyte sources on a schedule you control.
Incremental syncs
Pull only the records that changed since the last run instead of reloading everything.
One authorization
Authenticate GCS Data Lake once and Airbyte keeps every scheduled sync running on it.
Cloud or self-hosted
Run the GCS Data Lake connector on Cloud, Self-Managed Enterprise.
Sync capabilities
- Full Refresh SyncSupported
- Incremental SyncSupported
- Available onCloud, Self-Managed Enterprise
- Sources600+ Airbyte connectors
- Connector version1.0.10
What you'll need
- GCS Bucket NameThe name of the GCS bucket that will host the Iceberg data.
- Service Account JSONThe contents of the JSON service account key file. See the Google Cloud documentation for more information on how to obtain this.
- GCP LocationThe GCP location (region) for BigLake metastore resources. For example: "us-central1" or "us". See BigLake locations for available regions.
- Warehouse LocationThe root location of the data warehouse used by the Iceberg catalog. Must include the storage protocol "gs://" for Google Cloud Storage. For example: "gs://your-bucket/path/to/warehouse/
- Main Branch NameThe primary or default branch name in the catalog. Most query engines will use "main" by default. See Iceberg documentation for more information.
- Default NamespaceThe default namespace to use for tables. This will ONLY be used if the Destination Namespace setting is set to Destination-defined or Source-defined
- Catalog TypeSpecifies the type of Iceberg catalog (BigLake or Polaris).
Authenticate GCS Data Lake once
Service Account JSON
The contents of the JSON service account key file. See the Google Cloud documentation for more information on how to obtain this.
Related connectors
Common questions
What is ETL?
ETL, an acronym for Extract, Transform, Load, is a vital data integration process. It involves extracting data from diverse sources, transforming it into a usable format, and loading it into a database, data warehouse or data lake. This process enables meaningful data analysis, enhancing business intelligence.
What data can you extract from GCS Data Lake?
GCS Data Lake provides access to a wide range of data types, including: Structured data (organized into tables with defined columns and data types, such as CSV, JSON, and Avro files); Semi-structured data (some structure, but not necessarily a fixed schema, such as XML and JSON files); Unstructured data (no predefined structure, such as text, images, and videos); Time-series data (organized by time, such as stock prices, weather data, and sensor readings); Geospatial data (related to geographic locations, such as maps, GPS coordinates, and spatial databases); Machine learning data (used to train machine learning models, such as labeled datasets and feature vectors); and Streaming data (generated in real-time, such as social media feeds, IoT sensor data, and log files). Overall, GCS Data Lake's API provides access to a wide range of data types, making it a powerful tool for data analysis and machine learning.
How do I transfer data from GCS Data Lake?
This can be done by building a data pipeline manually, usually a Python script (you can leverage a tool such as Apache Airflow for this). This process can take more than a full week of development. Or it can be done in minutes on Airbyte in three easy steps: 1. Set up GCS Data Lake as a source connector (using Auth, or usually an API key). 2. Choose a destination (more than 50 available destination databases, data warehouses or lakes) to sync data to and set it up as a destination connector. 3. Define which data you want to transfer from GCS Data Lake and how frequently.
What are top ETL tools to transfer data from GCS Data Lake?
The most prominent ETL tools to transfer data to include: Airbyte, Fivetran, StitchData, Matillion, Talend Data Integration. These tools help in extracting data from various sources (APIs, databases, and more), transforming it efficiently, and loading it into and other databases, data warehouses and data lakes, enhancing data management capabilities.
What is ELT?
ELT, standing for Extract, Load, Transform, is a modern take on the traditional ETL data integration process. In ELT, data is first extracted from various sources, loaded directly into a data warehouse, and then transformed. This approach enhances data processing speed, analytical flexibility and autonomy.
Difference between ETL and ELT?
ETL and ELT are critical data integration strategies with key differences. ETL (Extract, Transform, Load) transforms data before loading, ideal for structured data. In contrast, ELT (Extract, Load, Transform) loads data before transformation, perfect for processing large, diverse data sets in modern data warehouses. ELT is becoming the new standard as it offers a lot more flexibility and autonomy to data analysts.
Start moving GCS Data Lake data today
Free for 14 days on Airbyte Cloud. Set up the GCS Data Lake connector once and let Airbyte keep it in sync.