Kafka to TiDB: How to Move Your Data

Move Kafka into TiDB with Airbyte. Why TiDB Cloud needs a TLS parameter, what the three-column output means, and whether Avro changes anything.

Summarize with AI:

Moving Kafka into TiDB gives events a durable home in a database that handles both transactional and analytical work. Topics expire on a retention policy chosen for operational reasons, so the messages your services exchanged last quarter are usually gone by the time anybody wants to query them.

This guide covers the managed path with Airbyte. Two things shape the build: what lands in TiDB is three columns rather than a modelled table, and whether your topics carry JSON or Avro changes what the pipeline can guarantee.

Kafka to TiDB at a glance:

CapabilitySupportedWhat it means for this pipeline
Output shapeThree columnsAn identifier, a timestamp and the payload as JSON
IdentifiersForced lowercaseTable, schema and column names are all lowercased
Database and schemaThe same thingTiDB does not distinguish between them
TiDB Cloud with TLSNeeds a JDBC parameterThe TLS protocol has to be named explicitly
Message formatJSON or AvroAvro needs a schema registry and a subject name strategy

Why move data from Kafka to TiDB?

Two situations account for most of these pipelines.

The first is retention with query access. Topics are sized for consumers that read promptly rather than analysts asking about last year, and TiDB keeps events indefinitely while remaining queryable with ordinary SQL that your existing tools already speak.

The second is keeping events beside operational data in a database that serves both workloads. If you want deep analytical modelling over event payloads rather than a queryable archive, Kafka to Snowflake handles semi-structured records more comfortably.

What do you need before you start?

Four things, and the second catches out anybody using the managed service:

A TiDB user with the right permissions. Create, insert, select, drop, create view and alter are all needed. The TiDB destination documentation lists them and the version requirements.

The TLS protocol named explicitly, on TiDB Cloud. Connecting to TiDB Cloud with TLS enabled requires specifying the protocol version in the JDBC parameters, which is not something the connection screen hints at.

Kafka bootstrap servers, read permission and existing topics. The connector consumes from topics that already exist, joins as a consumer group, and needs its own group identifier rather than one something else is using. The Kafka source documentation covers the settings.

A schema registry, if your topics carry Avro. Plus credentials if it is secured and a subject name strategy, so the right schema is chosen when deserialising.

If your cluster or database restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow lists before you begin.

How do you build a Kafka to TiDB pipeline in Airbyte?

Step 1: Settle the TLS parameter if you are on TiDB Cloud

Add the TLS protocol to your JDBC parameters before anything else, naming a specific version rather than assuming encryption negotiates itself. This is the connection problem people spend an afternoon on, because everything else about the configuration looks correct and the failure says nothing about protocol versions. Two minutes here saves that entirely.

Step 2: Configure the Kafka source

Click Sources in the left navigation, then New Source, and select Kafka, following adding a source. Supply the bootstrap servers, protocol, message format, subscription method and a dedicated group identifier. If your topics carry Avro, this is where the registry details go.

Step 3: Configure the TiDB destination

Click Destinations, then New Destination, and select TiDB, following adding a destination. Supply the host, port, database and credentials. Remember that TiDB treats a database and a schema as the same thing, and that identifiers are forced to lowercase, so name things in a way that survives that.

Step 4: Create the connection and sync inside retention

Click Connections, then New connection, select your streams and a sync mode. The schedule must sit comfortably inside your topic retention, because messages that expire before the pipeline reads them are lost rather than delayed. Alert on failure for the same reason.

Then build views over the JSON column, because nobody wants to write path expressions in every query.

What actually lands in TiDB?

Three columns per stream, and only one of them is yours. Each table gets an identifier assigned by Airbyte, a timestamp recording when the record was pulled, and a JSON column holding the event itself. That is a deliberately simple shape and it means no modelling happens on the way in.

For Kafka specifically that produces a mild nesting problem. Messages already arrive wrapped in an envelope carrying the record alongside its identifier and stream name, so what sits in the JSON column is a structure within a structure. Neither layer is difficult to reach into, and knowing both exist saves a confusing first query.

So build views that extract the fields your queries use into proper columns, which TiDB's MySQL-compatible JSON functions make straightforward. Do it once rather than in every query, and mind the lowercasing while you are there, since table and column names are forced to lowercase and a view referring to mixed-case names will not find them.

Does JSON or Avro change anything here?

It changes what you can rely on, though not what the tables look like. JSON messages need no extra configuration and carry no contract, so a producer adding or renaming a field does so silently. Avro messages need a schema registry, credentials where it is secured, and a subject name strategy so the connector selects the right schema when deserialising.

The benefit of Avro is upstream of this pipeline rather than inside it. A registry means producers cannot change a payload shape without the change being registered and validated, which is precisely the discipline that keeps your extraction views working. Without it, the first sign of a producer change is usually a view returning nulls.

Either way the payload lands in a JSON column, so the format is not a decision you make for TiDB's sake. What it does decide is whether your extraction views have a contract behind them, and if your organisation already runs a registry it is worth using here rather than treating the topics as plain JSON because that configuration is shorter.

Frequently asked questions

I cannot connect to TiDB Cloud.

With TLS enabled you must name the TLS protocol version in the JDBC parameters. The failure does not mention this, which is why it costs people an afternoon.

Why are my table names lowercase?

The destination forces all identifiers to lowercase, covering tables, schemas and columns. Write views accordingly rather than assuming the case you configured survives.

Which schema should I create?

TiDB does not distinguish a database from a schema, so a database is effectively the schema where your tables live. One name covers both concepts.

Do I need a schema registry?

Only for Avro messages, where you also need a subject name strategy. JSON needs nothing extra and gives you no contract in return.

Can I do this without writing code?

The pipeline, yes. The views extracting fields from the JSON column are SQL, and they are what turns a faithful archive into something people will query.

Get your Kafka data into TiDB

Name the TLS protocol in your JDBC parameters if you are on TiDB Cloud, since that is the connection failure that explains itself least. Give the pipeline its own consumer group and sync well inside your topic retention. Then expect three columns with the payload in JSON, build extraction views over it once rather than per query, and remember that identifiers arrive lowercased whatever you typed.

Airbyte's connector catalog includes 600+ pre-built connectors, so event streams can outlive the topics that carried them. For the same source into another operational database, see Kafka to PostgreSQL, and for a relational source into the same destination, PostgreSQL to TiDB.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.