Twitter to Kafka: How to Move Your Data

Move Twitter into Kafka with Airbyte. Why the search reaches back only seven days, and why the query defining your topic is invisible to every consumer.

Summarize with AI:

Moving Twitter into Kafka publishes matching posts where several systems can react to them. A sentiment model, an alerting service and an archive all want the same stream, and pointing each at the API separately means three sets of credentials competing for the same allowance.

This guide covers the managed path with Airbyte. Two things shape the build: the search reaches back only seven days, so anything you fail to capture is gone, and the query defining your stream is invisible to everybody consuming it.

Twitter to Kafka at a glance:

CapabilitySupportedWhat it means for this pipeline
HistorySeven daysThe start date cannot reach further back than that
End dateTen seconds priorIt must sit at least that far before the request time
Search queryDefines the datasetNothing in the messages records what it was
AuthenticationApp only bearer tokenRather than a user-authorised credential
DeliveryScheduled pollingThis is not a live firehose

Why move data from Twitter to Kafka?

Two situations account for most of these pipelines.

The first is fan-out to systems that treat posts differently. One consumer scores sentiment, another matches against a watchlist and raises alerts, a third keeps a record for compliance. Publishing once means one caller against a limited API rather than three teams independently consuming the same allowance.

The second is keeping anything at all, since the search window is short and a warehouse or archive downstream becomes the only durable record. With a single consumer, a bus is overhead and a direct pipeline into a database is simpler to build and easier to query afterwards.

What do you need before you start?

Four things, and the second one is the entire dataset definition:

An app only bearer token. Obtained from your developer account rather than by authorising as a user, which suits a pipeline. The Twitter source documentation covers the token and the date constraints.

A search query you have tested. This is not a filter over a dataset, it is the thing that decides what the dataset contains, and Twitter's query syntax rewards reading the guide before writing one.

Topics created in advance. The destination writes to topics that already exist. Give each distinct query its own topic rather than mixing them, for reasons the second half of this guide explains.

Somewhere durable for consumers to land it. Because the source keeps seven days, the bus is a transport rather than an archive, and something downstream has to be the long-term record.

If your organisation restricts access by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.

How do you build a Twitter to Kafka pipeline in Airbyte?

Step 1: Build and test the query before anything else

Write the query, run it against the API, and look at what comes back before you configure a source. A query that is too broad floods your consumers and burns your allowance; one that is too narrow quietly misses the posts you cared about, and with a seven-day window you cannot go back and check later. Twitter's query building guide is worth the read, since the syntax supports rather more than keyword matching.

Step 2: Configure the Twitter source

Click Sources in the left navigation, then New Source, and select Twitter, following adding a source. Supply the bearer token and your query. Both dates are optional and constrained: the start cannot be more than seven days ago, and the end must sit at least ten seconds before the request.

Step 3: Configure the Kafka destination

Click Destinations, then New Destination, and select Kafka, following adding a destination. Supply the bootstrap servers, security protocol and topic configuration. Messages are JSON, each wrapping the post alongside its identifier, the extraction timestamp and the stream name, so consumers read through an envelope.

Step 4: Create the connection and sync well inside the window

Click Connections, then New connection, select the stream and a sync mode. Schedule comfortably inside seven days and alert on failure, because a pipeline broken for longer than that has lost posts permanently rather than fallen behind.

Then record the query beside the topic name, because nothing in the messages themselves says what they were selected for.

Why is seven days the whole of it?

Because the connector reads the recent search endpoint, and recent means what it says. The start date cannot be set further back than seven days, so there is no backfill to run and no way to recover a period you were not watching. The end date has its own smaller constraint, needing to sit at least ten seconds before the request.

That makes a failed sync unusually expensive. On most pipelines an outage costs freshness and the data waits; here a pipeline down for a fortnight has permanently lost a week of posts, and no amount of investigation afterwards retrieves them. Alerting is not a refinement on this source, it is the thing that makes the archive trustworthy.

It also means the bus is transport rather than memory. Kafka retention is typically measured in days and the source offers seven, so unless something downstream is writing to durable storage, your organisation's record of what was said lasts exactly as long as whichever of those two windows is shorter. Decide which consumer owns the archive before anybody depends on it existing.

Why does the topic's meaning live outside the messages?

Because the search query decides membership and travels nowhere. A consumer receives posts that matched a query it cannot see, and nothing in the envelope or the record says which query that was. The topic means whatever somebody typed into a configuration screen, and that meaning is carried entirely by convention.

The dangerous version is a mid-stream change. Somebody broadens the query to include another brand term, the pipeline keeps running, and a consumer computing a sentiment average now averages a different population than it did yesterday without any signal that anything happened. The numbers move, the explanation is invisible, and the chart looks like a genuine shift.

So treat the query as part of the topic's contract. Document it beside the topic name, give distinct queries distinct topics rather than reusing one, and when a query genuinely needs to change, publish to a new topic rather than quietly altering the old one. That way a consumer's data means one thing for its whole life, which is what anybody computing a trend is implicitly assuming.

Frequently asked questions

Can I backfill older posts?

No. The start date cannot be more than seven days in the past, because the connector reads the recent search endpoint. Anything older was never available to collect.

Is this a live firehose?

No. Syncs are scheduled batches, so latency is your interval. Arriving on Kafka makes it look live, which is worth correcting before consumers are designed.

What happens if I change the query?

The topic's population changes with no signal to consumers. Publish to a new topic instead, so anything computing a trend is not silently comparing two different datasets.

Why does my end date get rejected?

It must be at least ten seconds before the request time, which catches people setting it to the current moment.

Can I do this without writing code?

The pipeline, yes. The consumers are yours, and one of them needs to be writing to durable storage if anybody expects a record beyond the retention windows.

Get your Twitter data into Kafka

Test the query before configuring anything, because it defines the dataset rather than filtering it, and a seven-day window means you cannot check later what you missed. Alert on failure, since an outage loses posts permanently. Decide which consumer owns durable storage. And treat the query as part of the topic's contract, documenting it and publishing to a new topic when it changes rather than altering the meaning underneath people.

Airbyte's connector catalog includes 600+ pre-built connectors, so public conversation can reach every system that watches it. For public publications onto the same bus, see Rss to Kafka, and for internal conversation onto the same bus, Slack to Kafka.

Start syncing now →

Integrate with 700+ apps using Airbyte

Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.