Slack to Databricks: How to Move Your Data
Move Slack into Databricks with Airbyte. Why thread replies dominate the sync cost, and why reconstructing conversations is the real work of the project.

Moving Slack into Databricks gives you somewhere to model how an organisation actually communicates. Slack search finds a message; it cannot tell you how long questions wait for answers, which channels carry real work, or whether a reorganisation changed anything.
This guide covers the managed path with Airbyte. Two things shape the build: fetching thread replies costs considerably more than fetching messages, and a conversation is not a row, so reconstructing one is modelling work rather than a query.
Slack to Databricks at a glance:
Why move data from Slack to Databricks?
Two situations account for most of these pipelines.
The first is analysing communication patterns, which needs derived measures rather than messages. Response times, thread participation and how work moves between channels are all computed from timestamps across several streams, and that is modelling work a lakehouse suits.
The second is keeping years of history, which a lakehouse does without a volume conversation. If the aim is finding messages rather than measuring them, a search engine is the better tool and Slack to Elasticsearch covers that route instead.
What do you need before you start?
Four things, and the first is a conversation with your workspace administrator:
A Slack app, and a decision about which channels it joins. The bot reads only channels it has been invited to, so the channel list is your dataset's boundary. The Slack source documentation covers the scopes and the setup.
A view on whether you need thread replies. They carry most of the substance in a working channel, and they are also where the request cost of this pipeline concentrates.
Permission to create Volumes in Unity Catalog. Staging goes through Avro files written into a Volume, which is separate from creating tables and worth requesting early.
Agreement about what this is for. Warehousing colleagues' messages deserves a deliberate decision and a stated purpose, because a dataset built for measuring response times can be read for other things entirely.
If your workspace restricts traffic by IP, add the Airbyte Cloud IP addresses to the allow list before you begin.
How do you build a Slack to Databricks pipeline in Airbyte?
Step 1: Decide which channels, and invite the bot to those
Pick the channels your analysis genuinely needs and add the bot to them explicitly, rather than reaching for the option that joins everything it can see. A narrower dataset is faster to sync, easier to justify and less likely to contain conversations nobody expected to be warehoused. Record the list somewhere, since the boundary exists only in Slack's membership and is invisible from the data.
Step 2: Configure the Slack source
Click Sources in the left navigation, then New Source, and select Slack, following adding a source. Supply your credentials and a start date, then select streams. Enable the option to skip threads with no replies, which removes a large number of pointless requests on any busy workspace.
Step 3: Configure the Databricks destination
Click Destinations, then New Destination, and select Databricks, following adding a destination. Supply the workspace details, warehouse or cluster, catalogue and schema. Message records carry nested attachments, reactions and block structures, and Spark reads those natively, so let them land intact.
Step 4: Create the connection and expect the first sync to be long
Click Connections, then New connection, select your streams and a sync mode. A workspace with years of history takes a while, and restarting a slow backfill discards progress without making anything faster. Daily is ample once caught up, since nobody needs a message modelled within minutes.
Then build the silver layer, because the landing tables describe Slack's objects rather than your organisation's conversations.
What does fetching threads actually cost?
Most of your request budget, because replies are not delivered alongside messages. Channel messages arrive in pages efficiently, and finding out whether a message has replies, then retrieving them, is separate work per message. On a workspace where most conversation happens in threads, that inverts where the time goes.
Skipping threads with no replies is the cheapest improvement available, since the great majority of messages in any channel are standalone and asking about their replies achieves nothing. It is one setting and on a large workspace it changes the character of the sync rather than trimming it.
Rate limiting deserves a mention alongside it, because Slack's response to exceeding limits is not uniform. Recovery behaviour differs between token types, and an application-level token that has been throttled can take considerably longer to return to normal throughput than a bot token would. The practical advice is the same either way: narrow the channel list, skip empty threads, and leave a slow backfill alone rather than restarting it.
How do you reconstruct a conversation?
By assembling it, because Slack does not hand you one. A conversation is a parent message plus its replies in order, spread across streams, and the thing people want to measure is the relationship between those rather than any individual row. That assembly is the actual work of this pipeline.
This is where the destination earns its place. Walking message hierarchies and computing gaps between timestamps is ordinary notebook work, and messages arrive with nested attachments and reaction structures that Spark reads without anybody flattening them first. The raw record sits beside the typed columns, so a field nobody modelled remains recoverable.
Build silver tables around the measures people asked for, typically one row per thread carrying the first response time, the number of participants and the duration. Then treat the definitions as the deliverable, because teams disagree about what counts as a response, and having that written into one modelled table prevents two dashboards reporting different numbers for the same week.
Frequently asked questions
Why is a channel missing?
The bot has not been invited to it. Channel membership is the dataset's boundary, and it is invisible from the data, so keep a record of which channels are covered.
Why is the first sync so slow?
Thread replies are fetched per message, so most of the time goes there. Skip threads with no replies and narrow the channel list before assuming something is broken.
Should I flatten message structures?
No. Spark reads nested attachments and reactions natively, so let them land and shape what you need in a silver table where the logic is visible.
Can I measure response times from one table?
No, it needs parent messages and their replies assembled together, which is why the raw events matter and why the modelling is the substance of this project.
Can I do this without writing code?
The pipeline, yes. Reconstructing conversations and agreeing what a response time means is modelling work, and it is where the value of this arrangement appears.
Get your Slack data into Databricks
Choose your channels deliberately and invite the bot to those, because membership is the boundary and nothing in the data reveals it. Skip threads with no replies, since replies are fetched per message and that is where the sync time goes. Then let nested message structures land intact and put the effort into silver tables that assemble conversations, carrying agreed definitions of what a response time actually is.
Airbyte's connector catalog includes 600+ pre-built connectors, so team communication can be measured alongside the work it supports. For the same source into a warehouse, see Slack to Snowflake, and for workspace documentation into the same destination, Notion to Databricks.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
