Hybrid Cloud Software for Data Movement Across Environments
Evaluate hybrid cloud software for data movement with evidence of where rows, credentials, logs, and compute run, and what crosses the control plane.

Hybrid data movement requires proof that every execution path respects your security boundary. This matters most when separate systems handle orchestration, execution, and storage. A deployment label tells you little, so identify where rows, credentials, logs, and compute run, then verify which network routes they use and which operators can access them. Treat sovereignty as an architectural property that you can test through configuration, network evidence, and repeatable controls.
TL;DR
- A hybrid-deployed tool and a hybrid data architecture are separate decisions, and a hybrid-deployed tool can still move regulated rows out of your boundary.
- Demand log-based CDC for every relational source and get the capture prerequisites in writing, including whether Oracle XStream capture needs a GoldenGate license.
- Rows staying in-boundary leave configuration, logs, and telemetry free to cross, so require a written field list with retention periods and storage regions.
- Score connector parity across deployment models and an inspectable exit path as sovereignty controls, and treat raw connector count as the least informative number a vendor shows you.
- Under DORA and the EU AI Act, pipeline log location and exit evidence belong in the contract as written requirements.
What Does Hybrid Cloud Data Movement Software Do?
Hybrid cloud software in this category copies changes from a source in one environment to a target in another, continuously, and tells you where the processing ran. That definition excludes hybrid storage and Integration Platform as a Service (iPaaS).
Hybrid deployment means the vendor runs orchestration and monitoring in its cloud, while an agent in your environment handles extraction and loading. A hybrid data architecture is a data-tier decision. Sensitive data stays in-boundary, while only anonymized or aggregated data moves to the cloud for analytics. A hybrid-deployed tool can still break a hybrid architecture if the agent in your network replicates a customer table to a cloud warehouse, because the rows leave the perimeter no matter where the control plane lives.
Which Movement Modes Should You Evaluate First?
Demand log-based CDC for every relational source, and treat trigger-based or query-based capture as a fallback when the database doesn't provide log access. Your ETL design for a hybrid ETL pipeline should already name which sources need CDC replication and which can run on a batch schedule.
Log-based CDC tails the database's redo or transaction log. It can capture inserts, updates, and deletes without querying production tables when retention, permissions, supported operations, and capture configuration are correct. Query-heavy batch ETL, by contrast, adds load to operational systems every time it runs.
The log is not free to open, and prerequisites vary across SQL databases. On Microsoft SQL Server, turning on CDC runs the sys.sp_cdc_enable_db and sys.sp_cdc_enable_table stored procedures, which need sysadmin and db_owner rights, respectively, and capture stops whenever SQL Server Agent is not running. On Oracle, XStream-based capture requires an Oracle GoldenGate license while LogMiner-based capture does not, so confirm which method each connector uses before you budget for it. Get those prerequisites and approval ownership in writing before your organization signs the order.
A continuous pipeline that must backfill existing rows typically starts with a snapshot, then streams changes from the log position recorded when the snapshot began. A pipeline that needs no historical backfill can start from a selected log position instead. Either way, the initial bulk load plus continuous CDC pattern keeps the source available with minimal performance impact, so budget the snapshot separately from steady-state latency. The capture method decides what touches the source; where the captured changes and their metadata travel next is a placement question.
How Do Control Plane and Data Plane Placement Change Sovereignty?
Splitting the control plane from the data plane can keep rows in your environment. The split also creates a second control plane metadata flow that buyers should ask vendors to itemize, so treat that metadata path as part of the system boundary.
In a hybrid control plane, the vendor hosts orchestration and monitoring while an agent in your network executes the sync. Databricks' architecture follows the same pattern, separating a control plane that holds the workspace application, notebooks, and configuration from a compute plane that processes data. Rows can stay in-boundary under this design while configuration, logs, and status telemetry still cross it, which is why the field list matters more than the architecture diagram.
A regional service boundary still does not answer who controls the software executing against the data. Ask whether you can inspect the code that runs inside your boundary and whether execution stays there. A repository verifies code access, and a network trace verifies execution location. If either check fails, require contractual controls for code-access rights and execution location, because every one of those flows also rides a network path with its own price.
What Do Network Paths and Egress Cost in a Hybrid Pipeline?
Provider pricing models put a cost on each path a replication stream can take. The total depends on direction, region, destination, monthly volume, and any free allowance.
Estimate your monthly change volume first, then apply the provider's current rate for the exact path you plan to run. Private-path support varies by service, source, and target, so check the support matrix for your source and target pair before you assume an inter-region transfer model applies.
A full deployment cost analysis includes compute, storage, orchestration, and monitoring, with egress as its own line. Cost only ranks the vendors that already clear your boundary requirements, and the scorecard is what tests those.
How Do You Score Software Against Hybrid Requirements?
Score capability evidence first and connector count last, because the count on a SaaS pricing page tells you nothing about what runs inside your boundary. Use written evidence for every score rather than relying on a product demonstration.
Score a criterion zero when the vendor cannot provide written evidence, and demand the filtered catalog as a contract exhibit. Schema evolution is the row a document cannot fully settle. Run hybrid integration testing against your own source and destination pair before production, and check whether the CDC service preserves column-to-value associations when a column is dropped from the middle of a table and whether it captures table truncations.
The downstream row matters because the same change stream should feed the warehouse, the lakehouse, and the retrieval stores your agents read. One governed pipeline keeps a single set of access controls in force for all of them, while a separate pipeline per consumer multiplies the boundary you have to audit. In regulated environments, that audit burden becomes a legal requirement.
How Do Regulations Change the Evidence You Need?
Regulated workloads require documented exit paths and auditable control plane exposure, so write both into the contract as architectural requirements.
The Digital Operational Resilience Act (DORA) has been in force since January 2025 and is now in active enforcement. The regulation requires financial entities to put exit strategies in place for information and communications technology (ICT) services supporting critical or important functions. A documented, inspectable exit path gives auditors evidence for DORA exit requirements.
The EU AI Act timeline moved in July 2026. The Digital Omnibus on AI defers high-risk obligations for stand-alone Annex III systems, including credit scoring and critical infrastructure, to December 2, 2027, while the Act's transparency obligations still apply from August 2, 2026. Treat pipeline log location as part of the audit scope for any high-risk AI system the pipeline feeds, and use the extra time to finish that review. Include the metadata field list in your processor agreement's scope so that hybrid logging and lineage across both planes is available as auditor evidence.
How Does Airbyte Flex Handle Data Movement Across Environments?
Airbyte Flex is a hybrid deployment built on a hybrid control plane. Airbyte operates the control plane while the data plane, along with your records, credentials, keys, and compute, runs in your environment, though some metadata such as cursor and primary-key values sits in the control plane. Orchestration, scheduling, upgrades, and monitoring run from that control plane, and Airbyte documents the outbound network requirements the data plane needs so your firewall team can review them.
Flex runs the same catalog of 700+ connectors as every other Airbyte deployment model, with no gating by deployment model, so the connector you evaluated is the connector that runs inside your perimeter. Log-based CDC sources read database logs and track log positions between syncs, and organization-level audit logging on supported paid tiers records events such as connection, permission, and source changes. Airbyte moves 26 billion records daily for 7,000+ companies.
The same replicated records are what an in-boundary retrieval stack indexes, so agent reads inherit the boundary and access controls the pipeline already enforces.
Where Should You Start?
The vendors worth taking into procurement are the ones whose written evidence matches what a network trace and a code review show. Use the completed scorecard to decide which source and target pairs clear that bar.
Airbyte built Airbyte Flex with CDC replication, audit logging, and an open-source foundation you can inspect and run yourself as the exit path.
Get a demo to see how Airbyte Flex would move data and route traffic in your hybrid environment.
Frequently Asked Questions
Can You Run Log-Based CDC From a Read-Only Replica?
Support depends on the database and the connector, so confirm it in the connector's source documentation before you design around it. On SQL Server, the CDC capture job runs on the primary even when change data is read from an Always On readable secondary, so the primary still carries capture load. Where a connector exposes lag metrics, track snapshot lag and streaming lag from the replica as separate numbers.
How Should You Distinguish Bring Your Own Cloud From a Hybrid Agent?
Treat Bring Your Own Cloud (BYOC) as an account and access question: whether applications and data run in your cloud account, and whether the vendor can access raw data there. A hybrid agent adds a network question on top, covering what it sends to the vendor's control plane, whether it needs only outbound connections, and which firewall rules it requires.
Should You Use AWS DMS for Continuous Replication?
Use AWS Database Migration Service (DMS) when it supports your exact source-target pair and its restart behavior meets your recovery requirements. For a continuously running pipeline, test failure recovery, schema changes, large-object handling, and restart behavior against your own workload before you commit.
Who Should Own the Evidence Review Before Procurement?
Assign an approval owner to each evidence document: security for network direction and control plane exposure, data engineering for capture method and schema behavior, and legal for license terms and the exit path. Record every zero-scored criterion as an unresolved contract item against its owner, so nothing unevidenced reaches signature.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
