Cloud Repatriation: When Moving Data Back On-Prem Pays Off
Cloud repatriation pays off per workload, not per company. See the signals, the break-even model, and the sovereignty cases that justify moving in-boundary.

Workload shape and regulatory tier decide whether a repatriation project succeeds, often before you price a single server. Treat each workload as the unit of decision rather than applying a company-wide cloud strategy uniformly. Your cloud bill tells you where to look, but the bill alone never settles the question. A well-matched move lowers recurring costs or satisfies a placement obligation. A poorly chosen one replaces elastic capacity with excess hardware, added staffing, and years of hybrid operations. The decision is a per-workload placement call, not a company-wide cloud exit.
TL;DR
- One enterprise CIO survey's 83–86% measures intent; a server and storage workload survey's 8–9% measures plans for full workload repatriation.
- Four conditions earn a break-even model: predictable growth, high sustained utilization, material egress, or a sovereignty obligation.
- Most models omit hardware refresh, added ops headcount, and cloud spend that keeps running during procurement.
- A CDC stream on a repatriated source crosses a network boundary when its destination remains outside that source's boundary.
What Do the Cloud Repatriation Surveys Actually Measure?
The headline repatriation numbers measure different actions, so a single percentage rarely means what it appears to. Cloud repatriation is the relocation of selected workloads and datasets from public cloud to infrastructure your organization controls, whether a colocation cage, a private cloud, or an owned data center. It covers five moves, and only one is a full cloud exit. A full exit removes every workload from public cloud, while partial repatriation moves a subset and creates a hybrid deployment model.
Repatriation Includes More Than Full Exits
Workload-level repatriation moves one application with its data. Data-only repatriation moves datasets in-boundary while compute stays in cloud, and phased repatriation sequences these across budget cycles. The headline surveys blur those categories. One enterprise CIO survey found 83% of enterprise CIOs planned to repatriate "at least some" workloads, which a single application satisfies.
Survey Percentages Measure Different Actions
A server and storage workload analysis puts full workload repatriation plans at 8–9%, with most organizations instead repatriating specific elements such as production data, backup processes, or compute. The 8–9% measures repatriation plans, not completed exits, which require a separate measure.
Uptime Institute's Global Data Center Survey 2025 counts 45% of IT workloads in corporate data centers and cloud's share stable at 10%. That figure describes current workload distribution, so the measures together suggest movement that is broad, shallow, and rarely a full exit.
Survey Sources Need Separate Interpretation
The table separates each survey's action measure from its commissioning interest.
Vendor-commissioned surveys report figures from 50% to 93%, which is why the decision cannot rest on a headline number. The workload's own properties settle it.
Which Signals Show That Repatriation Might Pay Off for You?
Seven workload properties settle this, and you should give each candidate its own data plane placement call. Your strongest candidates combine stable demand, predictable growth, material data-movement costs, loose coupling, operational readiness, or a placement obligation.
Workload Shape and Operating Conditions Come First
Read the scorecard below through a few assessment nuances rather than as a simple checklist. High and consistent utilization is what favors owned hardware, because intermittent demand never amortizes the purchase. Egress is a candidate signal, but no universal percentage threshold applies across providers, discounts, and workload shapes, so compare the actual egress share against the workload's total cloud spend rather than against a fixed cutoff.
Coupling and readiness carry the most hidden cost. Loose coupling means open formats and a replatform estimate that stays small against the first year of savings, while deep dependence on a managed service can swamp any storage savings. Operational readiness assumes you already hold colocation space and a platform team. Regulatory tier is the one dimension that can override the economics on its own, when the dataset falls under DORA exit scope or a jurisdictional access constraint.
Bounded Case Studies Clarify the Conditions
37signals is the bounded example. After leaving public cloud, it cut its annual cloud spend from $3.2M to roughly $1.3M, spending about $600,000 on Dell servers for colocation racks it already had.
The move added no staff. Its existing racks and staffing are material boundary conditions. Those figures leave out hardware refresh, new operations roles, and power and cooling. Your model should include each omitted cost.
Dropbox's historical infrastructure build covered a different operating period and does not establish a current repatriation benchmark. Ahrefs' engineering account describes infrastructure that never fully ran on public cloud. Neither example is repatriation evidence.
Candidate Scores Determine the Next Step
Score each candidate workload on these seven dimensions before opening a cost model.
A workload for which most of the seven dimensions favor moving advances to the break-even model. A workload for which most dimensions favor staying remains in cloud and goes to FinOps.
How Do You Calculate the Repatriation Break-Even Point?
Model both sides with six line items each, then divide one-time transition cost by the monthly cost difference. This comparison shows whether expected savings recover your transition cost within the payback threshold your organization sets before modeling.
On the cloud side, count instance cost, storage, egress, platform tax, licensing, and FinOps engineering time. For the in-boundary side, count five-year hardware amortization, power at your facility's Power Usage Effectiveness (PUE), colocation footprint, staffing, licensing, and refresh. This is a standard TCO model; the ETL total cost methodology covers each line, with transition cost and payback on top.
The break-even calculation must also count cloud spend that continues during procurement, engineering to replace managed services, egress, capex, replatform work, downtime, tooling, and compliance. Leaving any of these out makes moving look cheaper by construction.
In-boundary staffing appears as named headcount, while cloud staffing hides in the hours your engineers spend on Terraform, tagging, and commitment purchases. Carry a staffing line on both sides of the model or on neither, or the comparison tilts toward moving before the first server is priced.
Hardware refresh and failed procurement are real line items too. Refresh comes due regardless, and the Dell fleet is expected to last "five, maybe even seven years", so reserve on the shorter end. Procurement can also fail outright and leave you paying for hardware you cannot use, so budget a failed-procurement contingency or the model favors moving by construction.
When Is Staying in the Cloud the Better Decision for You?
If your team has not rightsized instances and bought commitment discounts, you have no repatriation case yet. You should first recover avoidable cloud waste and then compare the baseline after rightsizing and commitment discounts with the cost of moving.
Flexera's 2026 State of the Cloud Report puts wasted IaaS and PaaS spend at 29%, reversing a five-year decline. Recovering that reduces much of the monthly cost difference before any hardware enters the comparison.
Bursty demand is the clearest counter-signal. Elasticity is a cloud advantage when workloads require intense compute for short periods and then idle, so training often stays in cloud or runs as an in-boundary baseline with cloud burst.
Deep coupling to managed services is the second counter-signal: Snowflake's elastic warehouses, Databricks Unity Catalog, and BigQuery serverless have no in-boundary equivalent without re-engineering. Weigh your existing cloud data warehouses against that rebuild cost before assuming storage pricing settles the question.
Uptime Institute recommends a minimum of one to two qualified operators on site at all times for Tier III and Tier IV facilities. If your team migrated fully and let in-boundary skills lapse, you face a substantial talent acquisition exercise.
How Do Sovereignty Requirements Change Your Decision?
Residency describes data location, while sovereignty describes who can compel access to it. Either residency or sovereignty requirements can force a placement decision when the applicable law, contract, policy, or jurisdiction imposes that constraint.
For GDPR purposes, a US-headquartered provider can hold data in a Frankfurt data center while that data remains legally reachable by US government demands under the CLOUD Act. DORA is in active enforcement for EU financial entities. Article 28 requires firms to assess Information and Communication Technology (ICT) concentration risk and exit a provider without undue disruption, and Article 30 requires contracts to carry exit and transition provisions.
Reading those articles into a plan is a DORA exit strategy exercise. For placement, the principle is that portability is an architecture decision you make years before any exit.
A workable split can keep datasets that your applicable laws, contracts, or policies require you to control in an in-boundary tier. Depending on your obligations, those datasets may include payment card data, electronic protected health information (ePHI), or financial records that fall inside DORA scope. Processed aggregates derived from them can then flow to cloud analytics when applicable laws, contracts, or policies permit that flow.
A sovereign lakehouse design can hold that split together, since a dataset promoted into the governed core stays there even when someone later builds a convenient pipeline outward. Your obligations are documented location, assessed concentration, and a tested exit.
What Happens to Your Pipelines After Data Moves Back?
A repatriated source changes your pipeline path whenever its destination or control components remain outside the source's boundary. You need to re-evaluate network capacity, latency, schema propagation, credentials, and monitoring against the resulting architecture.
Boundary Crossings Change Pipeline Constraints
Once a source database returns in-boundary, any CDC stream, schema-change handler, or warehouse load whose destination remains outside that boundary must cross a network boundary. Streams and loads whose sources and destinations move into the same boundary do not necessarily cross one.
Published cases commonly moved application compute while leaving analytics out of scope. This leaves a hybrid path between the repatriated application and the remaining cloud services. Other cases moved object storage without documenting the full analytics stack. You should not assume the entire data platform moved with the cited workload.
Log reads that traveled inside a cloud region may now traverse your WAN link, so CDC latency budgets written for in-region hops need re-measuring. After a move, the constraint may become link capacity on your side. Schema evolution can cross the boundary too: a column added in-boundary now propagates into a cloud target across two change windows.
Full-Stack Moves Expand the Scope
You then choose whether to move the rest of the data stack or operate a hybrid path. Moving the whole data stack in-boundary means replacing integration tooling, orchestration, and monitoring too.
Hybrid Control Planes Preserve Pipeline Definitions
A hybrid control plane keeps orchestration, scheduling, and monitoring in a managed layer while connectors, credentials, and compute run inside your boundary. A repatriated source can therefore keep its pipeline definitions.
37signals moved application compute first and left roughly 10 petabytes in S3 under a multi-year storage contract. You should stage your exit expecting hybrid ETL pipelines for years.
How Do You Choose Between Full Exit, Hybrid, and Staying Put?
Your screening score and payback horizon produce three outcomes. A workload for which most dimensions favor moving and whose payback falls within your stated threshold goes to a pilot. A steady-state core plus burst or managed-service-coupled components goes hybrid, while the rest stay in cloud and go through FinOps.
You should expect hybrid to be the common answer. Most organizations repatriate specific elements rather than entire estates. Run a hybrid deployment readiness check first. Confirm that you have colocation capacity, a platform team, network paths to your cloud targets, and an identity model spanning both. Then move the highest-scoring workload first and take its pipelines with it.
How Does Airbyte Flex Repatriate Data Without Rebuilding Your Pipelines?
Airbyte Flex implements this hybrid deployment and includes managed upgrades. Airbyte operates the control plane while the data plane, along with your records, credentials, keys, and compute, runs in your environment, though some metadata such as cursor and primary-key values sits in the control plane. When a source moves from a cloud VPC to your data center, a connector inside that boundary reads it.
The same 700+ connectors and CDC run in either location, so moving a source in-boundary becomes a placement change rather than a connector replacement. Airbyte's open-source foundation keeps connector logic inspectable and supports portability, which reduces the risk that repatriation replaces cloud lock-in with integration-platform lock-in. The same replicated records are what an in-boundary retrieval stack indexes, so agent reads inherit the boundary and access controls the pipeline already enforces.
Airbyte processes 26 billion records daily across customer deployments. This platform-level scale can inform your evaluation, but it does not replace the workload-specific cost model that decides each placement.
Where Should You Start?
Start by using regulatory tier and workload shape to decide placement, then build a break-even calculation only for the candidates that survive that screen. Most estates end up hybrid, so treat repatriation as a sequence of per-workload placement decisions rather than a single cloud exit.
Airbyte keeps that sequence portable. Airbyte Flex runs the same connectors and CDC whether a source sits in cloud or in your own data center, and its open-source foundation keeps your integration layer inspectable and free of new lock-in.
Get a demo to see how Airbyte Flex supports hybrid deployment for the workloads you select.
Frequently Asked Questions
How Much Does It Cost You to Move a Petabyte Out of a Hyperscaler?
The cost depends on the provider's per-GB rate and any negotiated discount. At AWS retail egress pricing of about $0.08 per GB, one petabyte would cost roughly $80,000 before discounts. Microsoft and several hyperscalers now waive egress fees for customers fully leaving their platforms.
Do Your AI Workloads Belong In-Boundary?
Training rarely does, while steady production workloads often can. A cluster idling between runs may never reach the utilization needed to justify owned hardware, while consistently utilized production capacity can.
Why Do Your Lift-and-Shift Workloads Get Repatriated Most Often?
Uptime Institute found that organizations had lifted and shifted most repatriated applications into public cloud, where the applications "can neither grow to meet demand, nor shrink to reduce costs." Re-architecture usually beats relocation.
Can You Reverse a Repatriation?
Yes, more often than case studies admit. One practitioner reports moving the same workloads into and out of cloud repeatedly as corporate budgeting alternated between capital and operating expenditure, so keep formats open and pipelines portable to make each placement revisable.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
