mirror of
https://github.com/supabase/supabase.git
synced 2026-10-05 09:25:06 +03:00
[bot] Sync from supabase/troubleshooting (#40147)
* Sync from supabase/troubleshooting * fix(troubleshooting metdata): invalid toml in titles * format(troubleshooting) * lint(troubleshooting): apply mdx lint fixes --------- Co-authored-by: github-docs-bot <github-docs-bot@supabase.com> Co-authored-by: Charis Lam <26616127+charislam@users.noreply.github.com>
This commit is contained in:
6 files changed
+202
No files matched your search
@@ -21,6 +21,20 @@ jobs:
|
||||
with:
|
||||
persist-credentials: true
|
||||
|
||||
- name: Install pnpm
|
||||
uses: pnpm/action-setup@a7487c7e89a18df4991f7f222e4898a00d66ddda # v4.1.0
|
||||
with:
|
||||
run_install: false
|
||||
|
||||
- name: Use Node.js
|
||||
uses: actions/setup-node@49933ea5288caeca8642d1e84afbd3f7d6820020 # v4.4.0
|
||||
with:
|
||||
node-version-file: '.nvmrc'
|
||||
cache: 'pnpm'
|
||||
|
||||
- name: Install deps
|
||||
run: pnpm install --frozen-lockfile
|
||||
|
||||
- name: Decode the GitHub App Private Key
|
||||
id: decode
|
||||
run: |
|
||||
@@ -58,6 +72,7 @@ jobs:
|
||||
fi
|
||||
git checkout -b $BRANCH_NAME
|
||||
rsync --archive --verbose --ignore-existing ./troubleshooting-upstream/guides/ ./apps/docs/content/troubleshooting/
|
||||
pnpm format
|
||||
git add apps/docs/content/troubleshooting/
|
||||
if git diff --quiet --cached; then
|
||||
echo "No changes to sync"
|
||||
|
||||
+85
@@ -0,0 +1,85 @@
|
||||
---
|
||||
title = "High CPU and Slow Queries with `ERROR: must be a superuser to terminate superuser process`"
|
||||
topics = [ "cli", "database", "storage" ]
|
||||
keywords = []
|
||||
---
|
||||
|
||||
When facing high CPU utilization, slow query performance, and an `ERROR: must be a superuser to terminate superuser process` message regarding an autovacuum, it indicates that a critical, non-terminable autovacuum operation is running on your Postgres database. This guide explains why this happens and what steps you can take.
|
||||
|
||||
### **Core Postgres concepts**
|
||||
|
||||
To understand this issue, it's essential to grasp a few core Postgres concepts:
|
||||
|
||||
**What is MVCC (Multi-Version Concurrency Control)?**
|
||||
Postgres uses MVCC, which allows multiple transactions to access the same data simultaneously without locking each other out. Instead of updating a row in place, Postgres creates a new version of the row whenever data is modified or deleted. The old version remains, accessible to other transactions that started before the change.
|
||||
|
||||
**What are Dead Tuples?**
|
||||
The old versions of rows, which are no longer visible to any active transactions, are called "dead tuples" (or dead rows). These dead tuples consume disk space and can degrade performance if not cleaned up.
|
||||
|
||||
**What is Autovacuum?**
|
||||
Autovacuum is a set of background processes in Postgres designed to automatically reclaim storage occupied by dead tuples and update statistics for the query planner. It runs automatically when the number of dead tuples in a table crosses a certain configurable threshold (e.g., a percentage of the table's total rows, such as 20% in some default configurations).
|
||||
|
||||
**What is Transaction ID (XID) Wraparound?**
|
||||
Every transaction in Postgres is assigned a unique Transaction ID (XID). These XIDs are 32-bit integers, meaning there's a finite number of them (around 4 billion). If the database continuously creates new transactions without old ones being "frozen" (marked as permanently visible), the XIDs can eventually "wrap around" – meaning new transactions will be assigned XIDs that are numerically smaller than very old, still-active transactions. This makes it impossible for the database to determine which rows are visible and which are not, leading to potential data corruption and rendering the database unusable.
|
||||
|
||||
To prevent this critical issue, Postgres initiates a special autovacuum operation: the **"wraparound prevention vacuum."**
|
||||
|
||||
### **Understanding the problem: A critical autovacuum**
|
||||
|
||||
When you encounter the `ERROR: must be a superuser to terminate superuser process` associated with an autovacuum marked "to prevent wraparound," it signifies that this mandatory, system-critical operation is underway.
|
||||
|
||||
**Symptoms and Cause:**
|
||||
|
||||
- **High Resource Utilization:** You'll typically observe high CPU utilization (e.g., approaching 100%) and elevated Disk I/O, as the autovacuum process scans and cleans millions of rows. Memory usually remains stable.
|
||||
- **Slow Queries:** With resources heavily consumed by the autovacuum, other database queries will become significantly slower, potentially leading to application downtime.
|
||||
- **Non-Terminable Process:** Due to its vital role in preventing data corruption, a wraparound prevention autovacuum cannot be terminated, even by a superuser. Attempting to stop it will result in the `ERROR: must be a superuser to terminate superuser process` message (or the database immediately restarting it). It _must_ be allowed to complete.
|
||||
|
||||
**Why is it running?**
|
||||
This situation often arises in large, high-write tables (e.g., `your_table`, which might be hundreds of GBs in size and contain hundreds of millions of rows) that accumulate dead tuples rapidly. When the transaction ID age of the table approaches a critical threshold, Postgres automatically triggers this emergency autovacuum. For instance, if a table has millions of rows and over a million dead rows, exceeding its configured `autovacuum_vacuum_scale_factor` (e.g., 0.2), a regular autovacuum might initiate. However, if the XID age continues to increase, the system prioritizes the wraparound prevention vacuum to safeguard data integrity.
|
||||
|
||||
### **Mitigating performance impact during a critical autovacuum**
|
||||
|
||||
Since the wraparound prevention autovacuum cannot be stopped, the best approach is to provide the database with sufficient resources to complete the operation as quickly and efficiently as possible.
|
||||
|
||||
1. **Upgrade your Database Compute Instance:**
|
||||
|
||||
- **Action:** Temporarily scale up your instance's CPU (e.g., from `m6g.4xlarge` to `m6g.8xlarge` or higher).
|
||||
- **Why it helps:** More CPU cores and processing power will help the autovacuum operation run faster, reducing the overall time it impacts your database.
|
||||
- **Considerations:** This usually causes a brief downtime (typically 1-2 minutes) as the instance restarts. However, the autovacuum process is designed to pause and resume automatically.
|
||||
|
||||
2. **Increase Disk Throughput/IOPS:**
|
||||
- **Action:** If disk I/O utilization is also consistently high (e.g., near 100%), consider temporarily increasing your disk's provisioned IOPS and throughput.
|
||||
- **Why it helps:** Autovacuum is an I/O-intensive operation, involving a lot of reading and writing. Higher disk performance can significantly speed up the process.
|
||||
- **Considerations:** Cloud providers often have limitations, such as a cooldown period (e.g., 6 hours) between disk modification operations.
|
||||
|
||||
### **Monitoring progress and future prevention**
|
||||
|
||||
**Monitoring the Current Autovacuum:**
|
||||
You can monitor the progress of the active autovacuum processes using the `pg_stat_progress_vacuum` view:
|
||||
|
||||
```sql
|
||||
SELECT relid::regclass AS table, round(100.0 * heap_blks_scanned / heap_blks_total, 2) AS pct_scanned FROM pg_stat_progress_vacuum;
|
||||
```
|
||||
|
||||
This query will show the percentage of the table that has been scanned by the vacuum process. Once the `pct_scanned` reaches 100% for the critical table, the operation is largely complete, and resource usage should normalize. After it finishes, you can consider downgrading your instance and disk resources back to their original configuration.
|
||||
|
||||
**Preventive Measures for the Future:**
|
||||
To avoid future emergency wraparound vacuums, especially on high-read/write tables:
|
||||
|
||||
- **Monitor Dead Rows:** Regularly check the number of dead rows in your tables. You can use the following SQL query:
|
||||
|
||||
```sql
|
||||
select relname, n_live_tup, n_dead_tup, last_autovacuum
|
||||
from pg_stat_all_tables
|
||||
where schemaname = 'public'
|
||||
order by n_dead_tup desc
|
||||
limit 10;
|
||||
```
|
||||
|
||||
Or, if using Supabase CLI, you can run: `supabase db inspect bloat`
|
||||
|
||||
If a table consistently shows a high number of dead tuples (e.g., hundreds of millions of live rows with still a significant number of dead tuples even after a vacuum), it's a good indicator that proactive maintenance is needed.
|
||||
|
||||
- **Schedule Manual Vacuums:** For tables with heavy write activity, consider scheduling manual `VACUUM` or `VACUUM FULL` operations during off-peak hours to reclaim space and prevent XID age from becoming critical. `VACUUM FULL` is more aggressive but locks the table and rewrites the entire table, so it should be used with caution and planned downtime.
|
||||
|
||||
- **Adjust Autovacuum Settings:** For persistently problematic tables, your database team may need to adjust autovacuum parameters (like `autovacuum_vacuum_scale_factor`, `autovacuum_vacuum_threshold`, or `autovacuum_freeze_max_age`) to make autovacuum more aggressive or trigger earlier, preventing XID age from reaching critical levels.
|
||||
+25
@@ -0,0 +1,25 @@
|
||||
---
|
||||
title = "PGRST106: \"The schema must be one of the following...\" error when querying an exposed schema"
|
||||
topics = [ "auth", "database" ]
|
||||
keywords = []
|
||||
[[errors]]
|
||||
code = "PGRST106"
|
||||
message = "The schema must be one of the following: public, ..."
|
||||
|
||||
---
|
||||
|
||||
You may encounter a `PGRST106` error, stating `{"code":"PGRST106","message":"The schema must be one of the following: public"}`, when attempting to query a schema via the PostgREST API that you've recently exposed.
|
||||
|
||||
**Why This Happens:**
|
||||
This error occurs because the `authenticator` role's `pgrst.db_schemas` setting does not include the desired schema. This creates a conflict, as PostgREST relies on this setting to know which schemas it should expose.
|
||||
|
||||
**How to Resolve This:**
|
||||
To fix this, you need to update the `pgrst.db_schemas` setting for the `authenticator` role. Connect to your database and execute one of the following SQL commands:
|
||||
|
||||
- **To explicitly add your schema:**
|
||||
If you want to add your specific schema (e.g., `your_schema_name`) alongside existing schemas (like `public`), use the following command. Remember to replace `your_schema_name` with your actual schema name and include all other necessary schemas already defined.
|
||||
`ALTER ROLE authenticator SET pgrst.db_schemas = 'public, your_schema_name';`
|
||||
|
||||
- **To reset to [dashboard configuration](/dashboard/project/_/settings/api):**
|
||||
Alternatively, you can reset this setting to allow the Supabase dashboard's configuration to take effect automatically.
|
||||
`ALTER ROLE authenticator RESET pgrst.db_schemas;`
|
||||
@@ -0,0 +1,17 @@
|
||||
---
|
||||
title = "Realtime connections giving `TIMED_OUT` errors"
|
||||
topics = [ "realtime" ]
|
||||
keywords = []
|
||||
---
|
||||
|
||||
If your Realtime connections in your application are giving `TIMED_OUT` errors, this often indicates an incompatibility between the version of `realtime-js` being used in your `supabase-js` client library package and your Node.js version. This issue happens when using versions of Node.js older than v22 with more recent versions of `supabase-js`.
|
||||
|
||||
**Why This Occurs:**
|
||||
This happens due to a fundamental change made to fix various issues. See more details on why this change was made and solutions to fix it: https://github.com/orgs/supabase/discussions/37869
|
||||
|
||||
**How to Resolve This:**
|
||||
|
||||
There are two primary options to fix this:
|
||||
|
||||
1. **Upgrade Node.js**: Upgrade your Node.js version to the latest Long Term Support (LTS) release, such as Node.js v24 as of November 2025.
|
||||
2. **Explicitly set the WebSocket transport**: See the linked GitHub discussion above for a detailed guide on doing this.
|
||||
+26
@@ -0,0 +1,26 @@
|
||||
---
|
||||
title = "SSO Error: \"You do not have permissions to join this organization\" or prompts to create new organization"
|
||||
topics = [ "platform" ]
|
||||
keywords = []
|
||||
---
|
||||
|
||||
When attempting to log in via SSO, you may observe the message "You do not have permissions to join this organization" or be prompted to create a new organization.
|
||||
|
||||
**Why This Happens:**
|
||||
|
||||
Supabase treats email/password and SSO identities as distinct, even when the email addresses are identical. Existing memberships associated with email/password accounts are not automatically linked to a newly created SSO identity.
|
||||
|
||||
**How to Resolve This Issue:**
|
||||
|
||||
If you want your org members to auto-join:
|
||||
|
||||
1. As an organization administrator, ensure "Join organization on signup" is enabled in your SSO configuration.
|
||||
2. Navigate to your organization's `Members` settings.
|
||||
3. Remove any existing email/password-based user accounts that are intended to transition to SSO.
|
||||
4. Instruct these users to log in via SSO; they will be automatically re-added to your organization with the default role if auto-join is active.
|
||||
|
||||
Alternatively, you can manually re-invite their Google SSO identity to the org, ask them to accept the invite while signed in via Google and once confirmed, remove their old email/password account.
|
||||
|
||||
**Best Practice When Setting up SSO:**
|
||||
|
||||
Maintain owner account with a password login as a break-glass backup.
|
||||
+34
@@ -0,0 +1,34 @@
|
||||
---
|
||||
title = "Supabase CLI: \"failed SASL auth\" or \"invalid SCRAM server-final-message\""
|
||||
topics = [ "auth", "cli", "database", "supavisor" ]
|
||||
keywords = []
|
||||
---
|
||||
|
||||
When executing `supabase db push` or `supabase link` or any other authenticated actions from the Supabase CLI, you might encounter an authentication error with messages such as `failed SASL auth (invalid SCRAM server-final-message received from server)`.
|
||||
|
||||
**Why This Occurs:**
|
||||
This typically indicates an authentication failure where the database connection pooler (Supavisor) in certain scenarios may be incorrectly caching credentials for the internal Supabase role `cli_login_postgres` used for password-less flows with the CLI. This can lead to the your IP being temporarily banned from repeated failed attempts to connect.
|
||||
|
||||
**To resolve this, consider one of the following solutions:**
|
||||
|
||||
1. **Check Network Bans:**
|
||||
|
||||
- Navigate to your project's [Database Settings](/dashboard/project/_/settings/database) page.
|
||||
- Review any listed IP addresses that are blocked. Remove any entries that correspond to your current connection and then try the CLI action again.
|
||||
|
||||
2. **Use the old Password-Based authentication flow instead:**
|
||||
|
||||
- Provide your database password directly through an environment variable when running the CLI command.
|
||||
|
||||
```bash
|
||||
SUPABASE_DB_PASSWORD=<your-database-password> supabase db push
|
||||
```
|
||||
|
||||
3. **Skip the Pooler and connect directly to the database with the Supabase CLI (Requires IPv6):**
|
||||
|
||||
- If your network supports IPv6, you can use the beta CLI version with the `--skip-pooler` flag to bypass the connection pooler to avoid this particular issue.
|
||||
|
||||
```bash
|
||||
npx supabase@beta link --skip-pooler
|
||||
npx supabase@beta db push
|
||||
```
|
||||
Reference in new issue
Block a user