Confluence to BigQuery: How to Move Your Data
Move Confluence into BigQuery with Airbyte. Why the audit stream depends on your Atlassian plan, and why page bodies do not belong in the table.

Moving Confluence into BigQuery gives you a measurable view of a wiki that has grown past anybody's ability to survey it. How much documentation exists, who maintains it, which spaces have gone quiet: easy questions that Confluence answers slowly once an instance is large.
This guide covers the managed path with Airbyte. Two things shape the build: one stream depends on what Atlassian plan you are on and goes quiet if you are not, and page bodies are most of your volume and least of your value.
Confluence to BigQuery at a glance:
Why move data from Confluence to BigQuery?
Two situations account for most of these pipelines.
The first is documentation governance, meaning counting what exists, finding what has gone stale and seeing which teams maintain theirs. Those are aggregations over a few thousand rows and a warehouse answers them instantly.
The second is joining wiki activity to other engineering data you already hold here. If you intend to process the page content itself, Confluence to Databricks is the better home for that kind of work.
What do you need before you start?
Four things, and the second is worth checking before anybody promises a report:
An Atlassian API token, your domain and the account email. Authentication uses the email and token together, so both belong to the same account. The Confluence source documentation lists the fields and streams.
Knowledge of your Atlassian plan. The audit stream requires Standard or Premium, and on a lesser plan it will not be there, which matters because audit records are what a governance project usually wanted.
A BigQuery dataset in the right location. Location is fixed at creation and BigQuery will not join across locations, so put this where your other engineering data already lives.
A decision about page bodies. They are the difference between a small metadata table and one that costs money to query, and most governance questions do not need them.
If your network restricts traffic by IP, add the Airbyte Cloud IP addresses to the relevant allow list before you begin.
How do you build a Confluence to BigQuery pipeline in Airbyte?
Step 1: Check your plan before promising audit data
Find out which Atlassian subscription your Confluence instance is on, because the audit stream needs Standard or Premium and is the one most governance projects were actually after. Establishing that first avoids designing a report around records your plan does not expose, which is a conversation with a procurement outcome rather than a technical fix.
Step 2: Configure the Confluence source
Click Sources in the left navigation, then New Source, and select Confluence, following adding a source. Supply the API token, domain name and account email. Check which streams appear afterwards rather than assuming all five, since availability depends on both your plan and the account's permissions.
Step 3: Configure the BigQuery destination
Click Destinations, then New Destination, and select BigQuery, following adding a destination. Supply the project identifier, dataset and service account credentials. Without page bodies this is a small dataset, so batched standard inserts are ample.
Step 4: Create the connection and compare the stream list
Click Connections, then New connection, select your streams and a sync mode. Daily is generous for a wiki. Then check what actually arrived against the five documented streams, because a missing one here is a subscription fact rather than a fault.
Note that this connector reads Confluence Cloud, so a Data Center instance needs a different approach entirely.
Why might a stream simply not appear?
Because what you can read depends on what you pay Atlassian. The audit stream requires a Standard or Premium plan, and the connector checks whether the account can reach it rather than failing the whole sync, which is considerate behaviour and quiet behaviour at the same time.
The awkwardness is that audit records are usually the point. A project to understand who changed what, when permissions shifted or how a space was reorganised is asking for exactly that stream, and the other four describe what exists rather than what happened to it.
So separate the two kinds of question early. Pages, blog posts, spaces and groups answer how much documentation exists, where it lives and who can see it, which is a genuinely useful governance picture. Anything about change over time either needs the audit stream and the plan behind it, or needs you to accumulate your own snapshots and compare them, which this destination makes straightforward.
Should page bodies be in the table?
Usually not, for two reasons that compound. Page content arrives as storage and view representations, which are Confluence's markup rather than readable prose, so a column of it is neither pleasant to read nor simple to search without processing it first.
And it is almost all of your volume. The metadata describing a wiki of several thousand pages is a small table; the same wiki with every page body attached is a different proposition entirely, and in a warehouse billing on bytes scanned that turns an inexpensive governance dataset into one somebody queries carefully.
So take bodies only with a reason, and if you do, keep them out of the views people query by default. Governance questions need titles, authors, spaces and timestamps. Anything genuinely wanting the content, such as classifying documentation or finding duplicated guidance, is language work that belongs in a lakehouse where the markup can be parsed properly.
Frequently asked questions
Why is the audit stream missing?
It requires a Standard or Premium Atlassian plan, and the connector skips it rather than failing when the account cannot reach it.
Does this work with Confluence Data Center?
No. It is built against the Confluence Cloud REST API and configured with a Cloud domain, account email and API token.
Can I search page content from here?
Not comfortably. Bodies arrive as Confluence markup rather than plain text, so searching or classifying them is work for a destination built for that.
How do I see how documentation changed?
Either from the audit stream, if your plan includes it, or by accumulating snapshots and comparing extractions, which this destination handles cheaply.
Can I do this without writing code?
The pipeline, yes, and the governance questions need only ordinary SQL. It is the content questions that need a different destination.
Get your Confluence data into BigQuery
Check your Atlassian plan before designing anything, because the audit stream needs Standard or Premium and is usually what a governance project wanted. Compare the streams that arrived against the documented five. Then leave page bodies out unless you have a reason, since they are markup rather than prose and almost all of your volume in a warehouse that bills on bytes scanned.
Airbyte's connector catalog includes 600+ pre-built connectors, so a wiki can be measured without anybody reading it. For the same source into a warehouse with fine-grained controls, see Confluence to Snowflake, and for issue tracking into a lakehouse, Jira to Databricks.
Integrate with 700+ apps using Airbyte
Move data from 700+ sources into warehouses, lakes, and beyond. Set up pipelines in minutes with pre-built connectors and the Connector Builder.
