Files
supabase/apps/docs/content/guides/database/replication.mdx
T

101 lines
7.2 KiB
Plaintext

---
id: 'replication'
title: 'Database replication'
description: 'Compare read replicas, Supabase Pipelines, and manual replication.'
subtitle: 'An introduction to database replication and change data capture.'
sidebar_label: 'Overview'
---
Replication is the process of copying changes from your database to another location. It's also referred to as change data capture (CDC): capturing all the changes that occur to your data.
## Use cases
You might use database replication for:
- **Analytics and data warehousing**: Replicate your operational database to analytics platforms for complex analysis without impacting your application's performance.
- **Data integration**: Keep your data synchronized across different systems and services in your tech stack.
- **Backup and disaster recovery**: Maintain up-to-date copies of your data in different locations.
## Replication methods
Supabase supports three replication methods. Choose based on whether you need another Supabase Postgres database, a managed replication pipeline to a destination system, or full control over your own logical replication setup.
### Read replicas
Read replicas are additional Supabase Postgres databases kept in sync with your primary database. Use them when you want read-only query capacity, lower latency in another region, or to isolate analytical reads from application writes while staying inside Supabase Postgres.
- [Set up read replicas](/docs/guides/platform/read-replicas)
{/* supa-mdx-lint-disable-next-line Rule001HeadingCase */}
### Pipelines
<Admonition type="caution" title="Private Alpha">
Supabase Pipelines is currently in private alpha. Private alpha features can be unstable and may introduce breaking changes while we evaluate the product direction, refine the feature set, and incorporate customer feedback.
</Admonition>
Pipelines is a managed CDC feature for creating replication pipelines from Supabase Postgres to destination systems. It uses Postgres logical replication under the hood with the open-source [Supabase ETL engine](https://github.com/supabase/etl). A destination is where your replicated data is stored; the pipeline is the managed process that continuously streams database changes to that destination.
- [Set up Pipelines](/docs/guides/database/replication/pipelines)
#### Supported destinations
Pipelines currently supports BigQuery as the managed destination. We are working on new destinations, and this table will be updated as support expands.
| Destination | Insert | Update | Delete | Truncate | Schema change | Description |
| ------------------------------------------------------ | ------------ | ------------ | ------------ | ------------ | ------------- | ------------------------------------------------------------------- |
| [BigQuery](/docs/guides/database/replication/bigquery) | ✅ Supported | ✅ Supported | ✅ Supported | ✅ Supported | ✅ Supported | Managed replication to Google BigQuery for analytics and reporting. |
### Manual replication
Manual replication uses the same underlying Postgres logical replication features as Pipelines, but you configure and operate the pieces yourself. Use this path when you want to connect tools such as Airbyte, Estuary, Fivetran, Materialize, Stitch, AWS DMS, or another system that supports Postgres logical replication.
- [Set up manual replication](/docs/guides/database/replication/manual-replication-setup)
## Related features
For realtime features and syncing data to clients (browsers, mobile apps), see [Realtime](/docs/guides/realtime).
Realtime also uses Postgres changes, but it is intended for broadcasting database updates to clients rather than maintaining a copy of your database in another system.
## Concepts and terms
### Write-Ahead Log (WAL)
Postgres uses a system called the Write-Ahead Log (WAL) to manage changes to the database. As you make changes, they are appended to the WAL, which is a series of files (also called "segments") where the file size can be specified. Once one segment is full, Postgres will start appending to a new segment. After a period of time, a checkpoint occurs and Postgres synchronizes the WAL with your database. Once the checkpoint is complete, then the WAL files can be removed from disk and free up space.
### Logical replication and WAL
Logical replication is a method of replication where Postgres uses WAL files to transmit changes to another Postgres database, or to a system that supports reading WAL files.
### LSN
LSN is a Log Sequence Number that identifies a position in the WAL. It is often used to determine the progress of replication in subscribers and calculate the lag of a replication slot.
## Logical replication architecture
When setting up logical replication, three key components are involved:
- `publication` - A set of tables on your primary database that will be `published`
- `replication slot` - A slot used for replicating the data from a single publication. The slot, when created, will specify the output format of the changes
- `subscription` - A subscription is created from an external system (that is, another Postgres database) and must specify the name of the `publication`. If you do not specify a replication slot, one is automatically created
## Logical replication output format
Logical replication is typically output in two forms, `pgoutput` and `wal2json`. The output method is how Postgres sends changes to any active replication slot.
## Logical replication configuration
When using logical replication, Postgres keeps WAL files around for longer than it otherwise needs them. If the files are removed too soon, then your `replication slot` can become inactive or lost if the database receives a large number of changes in a short time.
In order to mitigate this, Postgres has many options and settings that can be [tweaked](/docs/guides/database/custom-postgres-config) to manage the WAL usage effectively. Not all of these settings are user configurable as they can impact the stability of your database. For those that are, these should be considered as advanced configuration and not changed without understanding that they can cause additional disk space and resources to be used, as well as incur additional costs.
| Setting | Description | User-facing | Default |
| ---------------------------------------------------------------------------------------- | ------------------------------------------------------ | ----------- | ------- |
| [`max_replication_slots`](https://postgresqlco.nf/doc/en/param/max_replication_slots/) | Max count of replication slots allowed | No | |
| [`wal_keep_size`](https://postgresqlco.nf/doc/en/param/wal_keep_size/) | Minimum size of WAL files to keep for replication | No | |
| [`max_slot_wal_keep_size`](https://postgresqlco.nf/doc/en/param/max_slot_wal_keep_size/) | Max WAL size that can be reserved by replication slots | No | |
| [`checkpoint_timeout`](https://postgresqlco.nf/doc/en/param/checkpoint_timeout/) | Max time between WAL checkpoints | No | |