Commercetools to Kafka: How to Move Your Data

Move commercetools into Kafka with Airbyte. Why scope errors appear at sync time, and why a repeated order is a version rather than a new event.

Summarize with AI:

Moving commercetools onto Kafka publishes commerce data where several systems can react to it. Fulfilment, fraud checking, customer messaging and analytics all want the same orders, and pointing four services at the same API means four sets of credentials competing for the same allowance.

This guide covers the managed path with Airbyte. Two things shape the build: permission problems surface when you sync rather than when you configure, and orders change state repeatedly, which is awkward on a bus that does not deduplicate.

Commercetools to Kafka at a glance:

CapabilitySupportedWhat it means for this pipeline
Stream listShows everythingRegardless of what your API client may actually read
Missing scopesError at syncLoud rather than silent, which is the better failure
Access neededRead onlyView scopes per resource, nothing that writes
Order stateChanges over timeSo the same order reaches the topic several times
DeliveryScheduled pollingThis is not an event stream from the platform

Why move data from Commercetools to Kafka?

Two situations account for most of these pipelines.

The first is fan-out to systems that treat orders differently. One service arranges fulfilment, another scores risk, a third updates a customer record, and publishing once means a single well-behaved caller against the API rather than several teams each building their own integration.

The second is feeding an existing event-driven architecture where Kafka is already the backbone. If only one system needs this data and the goal is analysis, a bus is overhead and a pipeline straight into a warehouse is simpler to build and far easier to query afterwards.

What do you need before you start?

Four things, and the first two come from the same screen:

Your project key and region. Both appear in the API URL shown under developer settings in the Merchant Center, and the region is one of a small set of specific values rather than a free-text field. The commercetools source documentation covers the setup.

An API client with read scopes for each resource you want. Created in the Merchant Center, and the secret is shown once at creation, so capture it then. Airbyte needs read-level access only.

Topics created in advance. The destination writes to topics that already exist, and giving orders, customers and payments their own topics is almost always better than combining them.

Agreement with your consumers about what a message means. Because an order arriving twice is normal here, and a consumer written on the assumption that it is not will behave badly.

If your cluster restricts inbound traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build a Commercetools to Kafka pipeline in Airbyte?

Step 1: Grant scopes for the resources you intend to sync

Decide which resources your consumers need, then create the API client with a view scope for each of them rather than granting broadly or minimally and adjusting later. The reason is timing: the connector will offer you every possible stream whatever your client can read, so the mismatch does not appear until a sync runs. Matching the scopes to the plan now means the first sync tells you something useful.

Step 2: Configure the commercetools source

Click Sources in the left navigation, then New Source, and select commercetools, following adding a source. Supply the project key, region, host, client identifier and secret, plus a start date. Both full refresh and incremental syncs are supported, and incremental is the sensible choice for a bus.

Step 3: Configure the Kafka destination

Click Destinations, then New Destination, and select Kafka, following adding a destination. Supply the bootstrap servers, security protocol and topic configuration. Messages are JSON, each wrapping the record alongside its identifier and the stream it came from, so consumers read through an envelope.

Step 4: Create the connection and read the first errors carefully

Click Connections, then New connection, select your streams and a sync mode. An error naming a resource is almost certainly a missing scope rather than a fault, so read it before investigating anything else. Commercetools applies rate limits, so a gentle schedule is also a considerate one.

Then tell your consumer teams how to treat repeated orders, because that assumption is the one that breaks things quietly.

Why do permission problems appear so late?

Because the connector shows you what commercetools offers rather than what your client can reach. The documentation is explicit that the interface lists all possible data sources and raises errors during syncing if it lacks permission for a resource, which means stream selection is a menu rather than an inventory.

That is worth knowing and it is also the better design. Plenty of connectors handle a missing permission by quietly omitting the stream, which produces a sync that succeeds and a dataset that is incomplete for reasons nobody can see. An error naming the resource you cannot read is a considerably kinder failure, even though it arrives later than you would like.

So treat the first sync as part of setup rather than the end of it, and read any resource-specific error as a scope question first. Grant read-level scopes only while you are there, since nothing this pipeline does requires the ability to change an order, and a client that can write to your commerce platform is a risk nobody chose deliberately.

What does a repeated order mean to a consumer?

That it is a version rather than an event, which is the distinction consumers most often get wrong. An order moves through states as it is paid, picked, shipped and sometimes returned, and each sync that catches it in a new state publishes it again. Four messages about one order is normal operation here.

A consumer counting messages therefore counts state changes rather than orders, which is a quietly wrong number that looks plausible. A fulfilment service acting on every message it sees may act more than once on the same order. Neither failure announces itself, and both come from treating a polled snapshot as though it were an event emitted at the moment something happened.

Compaction does not rescue you either, because messages are keyed by an identifier the pipeline generates rather than by the order, so the broker cannot tell that two messages describe the same thing. Consumers have to deduplicate on the order identifier themselves and decide what a state transition means to them, and that expectation belongs in writing before the first consumer is built rather than after the second one misbehaves.

Frequently asked questions

A stream errors naming a resource. What is wrong?

Almost certainly a missing scope on your API client. The interface lists every possible stream regardless of what your client can read, so the mismatch appears at sync time.

Where do I find my project key and region?

Both appear in the API URL under developer settings in the Merchant Center, and the region must be one of the specific supported values rather than a description.

Why does the same order appear several times?

Orders change state, and each sync catching a new state publishes the order again. Consumers should deduplicate on the order identifier rather than counting messages.

Can compaction collapse the repeats?

No, because messages are keyed by a generated identifier rather than by the order, so the broker cannot recognise two messages as the same thing.

Can I do this without writing code?

The pipeline, yes. The consumers are yours, and the deduplication logic they need is the part that makes this arrangement correct.

Get your Commercetools data into Kafka

Grant a read scope for each resource before the first sync, since the stream list shows everything and permission errors only appear when data moves. Take the project key and region from the API URL in the Merchant Center. Then write down, for your consumer teams, that an order arriving four times is four versions of one order rather than four orders, because compaction cannot fix it and a plausible wrong number is worse than an obvious one.

Airbyte's connector catalog includes 600+ pre-built connectors, so commerce data can reach every system that reacts to it. For customer records onto the same bus, see Salesforce to Kafka, and for another business platform onto the same bus, Microsoft Dataverse to Kafka.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.