mirror of
https://github.com/supabase/supabase.git
synced 2026-10-05 09:25:06 +03:00
feat(reports): add memory commitment chart to database report DEBUG-119 (#46435)
## Problem Sustained memory overcommitment is one of the failure patterns that most often leads to databases being killed when the system runs out of memory. Today the database report only shows used memory against total RAM, which hides the kernel's commit accounting: a project can sit far above its commit limit (RAM plus swap, adjusted by the overcommit ratio) and the dashboard gives no signal until something breaks. A combined chart with swap, overcommitment, and main memory was considered too dense, and a standalone swap chart was not useful enough on its own. Linear: [DEBUG-119](https://linear.app/supabase/issue/DEBUG-119) ## Fix Adds a separate "Memory commitment" chart between the existing memory usage and swap charts. It plots `ram_commit_used` (Committed_AS) as the main series with `ram_commit_limit` (CommitLimit) as the max-value threshold line, so values approaching or crossing the limit are visually obvious. Backend support for the two new metric attributes ships in supabase/platform#33321. ## How to test - Wait for the platform PR (supabase/platform#33321) to deploy so the two new attributes are accepted by the infra monitoring endpoint. - Open any project, navigate to Database -> Reports. - Confirm the new "Memory commitment" chart appears between "Memory usage" and "Swap usage". - Confirm the chart shows two series: the committed memory bars/area and a max line for the commit limit. - Hover the legend and the chart to confirm the tooltips read clearly. - Confirm the time range selector and chart sync (`syncId: 'database-reports'`) keep this chart aligned with the others. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Added comprehensive guide for interpreting memory commitment patterns in telemetry reports, including component breakdown, chart pattern guidance, and actionable recommendations. * **New Features** * Added memory commitment chart to database observability dashboard, displaying RAM commitment usage and limits. * Extended monitoring API to support new RAM commitment metrics. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
This commit is contained in:
1 parent
de26826f3c
commit
9024f02f25
3 files changed
+78
-3
No files matched your search
@@ -137,6 +137,41 @@ Actions you can take:
|
||||
| [Tune Postgres configuration](https://pgtune.leopard.in.ua) | Improve memory management settings |
|
||||
| Implement application caching | Add query result caching to reduce memory load |
|
||||
|
||||
### Memory commitment
|
||||
|
||||
The Memory commitment chart shows how much memory the Linux kernel has promised to processes (`Committed_AS`) against the maximum it is willing to promise (`CommitLimit`). It is a leading indicator of out-of-memory risk that is not visible on the Memory usage chart.
|
||||
|
||||
Memory is committed when a process asks the kernel for an allocation (for example via `malloc`, a stack growth, or `mmap`). The kernel records the promise immediately, but only assigns physical pages when a page is first written. Committed memory is therefore the sum of every outstanding promise across every process, regardless of whether those pages have been touched yet.
|
||||
|
||||
| Component | Description |
|
||||
| ---------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| **Committed** | Total memory the kernel has promised to processes (`Committed_AS`). Includes promises that have not yet been backed by physical pages. |
|
||||
| **Commit limit** | Maximum memory the kernel will commit (`CommitLimit`). Derived from physical RAM, swap, and the kernel's overcommit ratio. Acts as the danger threshold for this chart. |
|
||||
|
||||
How to read it:
|
||||
|
||||
| Pattern | What it means |
|
||||
| ------------------------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------- |
|
||||
| Flat Committed well below the Commit limit | Healthy. Connections and queries are sized for the compute tier. |
|
||||
| Committed gradually rising over days or weeks | Organic growth or a memory leak. Investigate connection counts, long-lived prepared statements, and extensions before usage outgrows the tier. |
|
||||
| Committed spiking near or above the Commit limit | Dangerous. Usually a connection storm or several large concurrent queries. The next spike may trigger the OOM killer and crash Postgres. |
|
||||
| Committed sustained above the Commit limit | The instance is on borrowed time. Plan an upgrade or fix the workload before the next out-of-memory event. |
|
||||
|
||||
Why this matters for Postgres:
|
||||
|
||||
- Each new connection is a `fork()` of the postmaster, which inflates Committed_AS by roughly the size of `shared_buffers` until copy-on-write pages diverge. Connection bursts can therefore blow past the Commit limit long before physical memory is exhausted.
|
||||
- Each query can allocate up to `work_mem` per sort or hash node. A handful of expensive concurrent queries can push commit far above what the Memory usage chart reports as "used".
|
||||
- When the kernel cannot honor its promises, the OOM killer terminates a process. On a database server that is usually a Postgres backend or, worse, the postmaster, which takes the whole database down.
|
||||
|
||||
Actions you can take:
|
||||
|
||||
| Action | Description |
|
||||
| --------------------------------------------------------------------------------------------------- | --------------------------------------------------------------------------------------- |
|
||||
| [Use connection pooling](/docs/guides/database/connecting-to-postgres#how-connection-pooling-works) | Route clients through Supavisor or PgBouncer to cap fork-driven commit pressure. |
|
||||
| Lower `max_connections` or per-pool sizes | Reduce the upper bound on concurrent backends so spikes cannot exceed the Commit limit. |
|
||||
| [Tune `work_mem`](https://pgtune.leopard.in.ua) | Reduce per-operation memory allocations on workloads with many concurrent queries. |
|
||||
| [Upgrade compute size](/docs/guides/platform/compute-and-disk#compute-size) | Raise both physical RAM and the Commit limit so the workload fits with headroom. |
|
||||
|
||||
### CPU usage
|
||||
|
||||
<Image
|
||||
|
||||
@@ -19,7 +19,8 @@ export const getReportAttributesV2: (
|
||||
maxConnections?: MaxConnectionsData,
|
||||
pgBouncerMaxConnections?: number,
|
||||
isSpendCapEnabled?: boolean,
|
||||
showDiskIOBurstBalanceChart?: boolean
|
||||
showDiskIOBurstBalanceChart?: boolean,
|
||||
showMemoryCommitmentChart?: boolean
|
||||
) => ReportAttributes[] = (
|
||||
entitledFeatures,
|
||||
project,
|
||||
@@ -27,7 +28,8 @@ export const getReportAttributesV2: (
|
||||
maxConnections,
|
||||
pgBouncerMaxConnections,
|
||||
isSpendCapEnabled,
|
||||
showDiskIOBurstBalanceChart
|
||||
showDiskIOBurstBalanceChart,
|
||||
showMemoryCommitmentChart
|
||||
) => {
|
||||
const computeVariantId = mapComputeSizeNameToAddonVariantId(project?.infra_compute_size)
|
||||
const provisionedDiskIops = diskConfig?.attributes?.iops
|
||||
@@ -92,6 +94,42 @@ export const getReportAttributesV2: (
|
||||
},
|
||||
],
|
||||
},
|
||||
{
|
||||
id: 'memory-commitment',
|
||||
label: 'Memory commitment',
|
||||
docsUrl: `${DOCS_URL}/guides/telemetry/reports#memory-commitment`,
|
||||
hide: !showMemoryCommitmentChart,
|
||||
showTooltip: true,
|
||||
showLegend: true,
|
||||
hideChartType: false,
|
||||
defaultChartStyle: 'bar',
|
||||
showMaxValue: true,
|
||||
showGrid: true,
|
||||
syncId: 'database-reports',
|
||||
valuePrecision: 2,
|
||||
YAxisProps: {
|
||||
width: 75,
|
||||
tickFormatter: (value: number) => formatBytesMinMB(value, 2),
|
||||
},
|
||||
attributes: [
|
||||
{
|
||||
attribute: 'ram_commit_used',
|
||||
provider: 'infra-monitoring',
|
||||
label: 'Committed',
|
||||
tooltip:
|
||||
'Total memory the kernel has promised to processes (RAM plus swap). Sustained values near or above the commit limit indicate overcommitment and a high risk of out-of-memory failures',
|
||||
},
|
||||
{
|
||||
attribute: 'ram_commit_limit',
|
||||
provider: 'infra-monitoring',
|
||||
label: 'Commit limit',
|
||||
isMaxValue: true,
|
||||
omitFromTotal: true,
|
||||
tooltip:
|
||||
'Maximum memory the kernel will commit (RAM plus swap, adjusted by the overcommit ratio). Committed memory approaching this limit puts the database at risk of being killed when the system runs out of memory',
|
||||
},
|
||||
],
|
||||
},
|
||||
{
|
||||
id: 'swap-usage',
|
||||
label: 'Swap usage',
|
||||
|
||||
@@ -137,6 +137,7 @@ const DatabaseUsage = () => {
|
||||
project?.cloud_provider !== 'FLY'
|
||||
|
||||
const showDiskIOBurstBalanceChart = useFlag('showDiskIOBurstBalanceChart')
|
||||
const showMemoryCommitmentChart = useFlag('showMemoryCommitmentChart')
|
||||
|
||||
const REPORT_ATTRIBUTES = getReportAttributesV2(
|
||||
entitledFeatures,
|
||||
@@ -145,7 +146,8 @@ const DatabaseUsage = () => {
|
||||
maxConnections,
|
||||
defaultMaxClientConn,
|
||||
isSpendCapEnabled,
|
||||
showDiskIOBurstBalanceChart
|
||||
showDiskIOBurstBalanceChart,
|
||||
showMemoryCommitmentChart
|
||||
)
|
||||
|
||||
const { isPending: isUpdatingDiskSize } = useProjectDiskResizeMutation({
|
||||
|
||||
Reference in new issue
Block a user