Files
supabase/apps/studio/components/interfaces/Settings/Infrastructure/InfrastructureConfiguration/HaTopology.utils.ts
T
Alaister YoungandAlaister Young 928049ce2c [FE-3717] feat(studio): Multigres cluster topology diagram (#49298)
Adds an infrastructure/topology diagram for High Availability
(Multigres) projects showing the real cluster topology — gateway tier,
shard group, and the primary + read replicas inside it — on both the
project homepage and the database/replication page, replacing the
primary-only view and the "Replication unavailable" empty state.

<img width="790" height="541" alt="Screenshot 2026-08-20 at 8 23 42 PM"
src="https://github.com/user-attachments/assets/0bce21e3-2091-4285-84ca-60fdecb10d39"
/>

Addresses
[FE-3717](https://linear.app/supabase/issue/FE-3717/show-replicas-in-replication-diagram).

**Added:**

- `data/ha-admin/` — read-only queries for the mgmt-api
`/ha-admin/v1/{gateways,poolers,cells,databases}` multiadmin passthrough
(ported from `bobbie/ha-stub`, re-authored to `queryOptions`). Responses
are validated with zod at the fetch boundary (all fields optional per
proto3 zero-value omission; enum-shaped fields stay plain strings so new
proto values degrade gracefully); malformed payloads surface through the
diagram's error fallback.
- `HaTopology.utils.ts` — pure topology mapper (+ 26 unit tests): shard
grouping, primary identified via `routingState.role` (deprecated `type`
as fallback) with **failover-safe election** — when the outgoing and
incoming primary briefly both claim `ROUTING_ROLE_PRIMARY`, the highest
routing rule (coordinator term, leader subterm) wins, matching the
multigateway's own election — plus status mapping onto the existing
Healthy / Coming up / Going down / Unhealthy vocabulary, and an AZ
formatter for `id.cell` that degrades to the raw cell name.
- HA diagram nodes/edges: `Multigateway` card, shard group box with
header pill (`Shard 1`, `Automatic failover` + tooltip), `Primary
Database` card styled like the standard diagram's — neutral border,
green icon chip (with the standard CPU / Disk / RAM footer — connections
omitted until their meaning through the multigateway is confirmed),
`Read Replica` cards, and the standard animated replication edges
(status lives on the card badges). Poolers and gateways poll every 30s
without re-running layout (topology projection + structural sharing).
Drag-to-pan works through the shard group box, and the metrics footer's
skeleton matches the loaded row height so the card doesn't shift.
- Accessibility: the failover tooltip trigger is a keyboard-focusable
button, status badges sit in stable `role="status"` live regions, the
region flag is decorative (`alt=""`), and the edge dash/spinner
animations respect `prefers-reduced-motion` (applied to the pipelines
diagram's edges too).
- Fallbacks: `AlertError` ("Failed to retrieve cluster topology") when
either ha-admin query errors, and a "Cluster topology unavailable" empty
state when the topology comes back empty — never a half-rendered
diagram.

**Changed:**

- `InstanceConfiguration` is now topology-source-aware: it branches
internally on `useHighAvailability()`, so both surfaces (homepage
`TopSection` and the replication page) get the right diagram with no new
wiring. The two-pass measured dagre layout moved into a shared
`DiagramFlow`; `nodeTypes`/`edgeTypes` are module-level consts.
- `getEdgeVisual` + the mid-edge icon chip lifted out of
`ReplicationDiagram/Edges.tsx` into
`components/ui/ReactFlow/EdgeVisual.tsx` so both diagrams derive edge
icon + line style from one state object (no behavior change for the
pipelines diagram). The primary card's CPU/Disk/RAM footer is likewise
extracted into a shared `ComputeMetricsFooter`.
- Fixes a latent relayout loop inherited from the region-box pattern:
handing React Flow a freshly created (unmeasured) group node on every
layout pass reset `nodesInitialized`, re-triggering the measured pass
and `fitView` forever — which made the diagram snap back to center and
effectively unpannable. The shared `DiagramFlow` now re-attaches known
measurements to group nodes, which also covers the standard diagram's
region boxes.
- Standard diagram: the API Load Balancer → primary edge is now static —
no data flows over it, the line only indicates a relation.
- `database/replication` page: the HA early-return empty state is
replaced by the diagram under a "High Availability cluster topology"
header. Non-HA projects are untouched.

**Intentional deviations from the mock** (for design review):

1. **No per-replica regions** — alpha replicas are one-per-cell inside a
single region, so the mock's `eu-west-1` / `ap-southeast-1` on sibling
replicas would be false. Availability zone per node, region shown once
on the primary.
2. **"Primary Database", not "Main Database"** — matches the string both
existing diagrams already ship, and the same component now renders both
project types.
3. **No collapse chevron on the shard header** — alpha has exactly one
shard; collapsing it would hide the whole diagram. The group box still
ships; add collapse when `shards.length > 1`.
4. **Failover shown on the shard group, not replica cards** — failover
is a cohort property; per-card badging would assert readiness we can't
verify without a per-pooler `/status` fanout.
5. **Standard node/edge styling reused** (per review) — neutral primary
border + green chip and the default animated edges instead of the mock's
green ring and dashed green arrowed edges, keeping the HA and non-HA
diagrams visually consistent.

**Confirmed against a real local Multigres cluster:** cells are named
`cell-1`/`cell-2`/… (not AZ-shaped — the AZ formatter falls back to the
raw cell name as designed); `GET /platform/projects/{ref}/databases`
returns only the primary row for HA projects; and the `/ha-admin`
passthrough returns **each gateway/pooler record once per cell it fans
out to** — the topology mapper dedupes by id, but worth confirming with
@sbc-bobbie whether the backend should dedupe.

**Known alpha limitation:** node health and the "replicating" edge state
derive from the pooler's *topology record*
(`lifecycleStatus`/`servingStatus`), not a live probe — a pooler that
crashes without publishing a terminal state can read as healthy until
the topology evicts its record, and a serving replica with paused replay
still shows a green edge. This matches the existing replication
diagram's semantics (`ACTIVE_HEALTHY` ⇒ animated edge). Live per-pooler
signals (WAL receiver state, replay position) exist on `GET
/poolers/{cell}/{name}/status` but need a per-pooler fanout —
deliberately deferred, noted on `getPoolerStatus`.

**Still to confirm** (doesn't block review): whether the `/ha-admin`
passthrough is deployed to production or staging-only (if staging-only,
this should get a flag before GA).

## To test

Tested end-to-end locally against a real Multigres project
(standard-project regression pass, HA creation flow, error fallback
against real 500s, and full topology + polling + console checks against
live multiadmin data):

- **HA project homepage**: diagram card shows Multigateway → shard box
(`Shard 1`, count badge, `Automatic failover` tooltip) → green-bordered
Primary Database (region, AZ, size) + Read Replica cards (AZ), dashed
green animated edges to healthy replicas. No flow/map toggle for HA.
- **HA project → Database → Replication**: same diagram under a "High
Availability cluster topology" header; no Destinations section; the old
"Replication unavailable…" state is gone.
- **Error path**: if `/ha-admin/v1/*` fails, both surfaces show "Failed
to retrieve cluster topology" with Contact support — no partial diagram.
- **Standard project regression**: homepage diagram (primary card, flow
⇄ map toggle round-trips), replication page (pipelines diagram +
Destinations) all unchanged; zero requests to `/ha-admin/*`.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
  - Added High Availability topology diagrams to the Replication page.
- Display gateways, primary databases, replicas, shards, statuses,
regions, infrastructure details, and compute metrics.
- Added observability links and live topology updates with loading,
error, and unavailable states.

- **Bug Fixes**
- Improved handling of incomplete infrastructure identities and
unexpected data.
  - Corrected topology layout, node spacing, and visual edge behavior.

- **Accessibility**
- Reduced-motion preferences now disable diagram animations and loading
effects.
  - Improved status announcements for assistive technologies.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Alaister Young <10985857+alaister@users.noreply.github.com>
2026-08-26 16:51:36 +08:00

193 lines
7.5 KiB
TypeScript

import { groupBy, partition, uniqBy } from 'lodash'
import type { Multigateway } from '@/data/ha-admin/ha-cluster-gateways-query'
import type { HaClusterPoolersData, Multipooler } from '@/data/ha-admin/ha-cluster-poolers-query'
/**
* Pure helpers mapping the multiadmin cluster state (gateways + poolers) onto
* the shapes the High Availability infrastructure diagram renders. Every field
* on the multiadmin responses is optional because proto3 JSON omits zero
* values — an absent field means "the default", not "missing data".
*/
export type HaPoolerStatus = 'healthy' | 'coming_up' | 'going_down' | 'unhealthy'
export interface HaShard {
id: string
name: string
primary?: Multipooler
replicas: Multipooler[]
}
export interface HaTopology {
gateways: Multigateway[]
shards: HaShard[]
}
export const getPoolerKey = (pooler: Pick<Multipooler, 'id'>) =>
`${pooler.id?.cell ?? 'unknown'}-${pooler.id?.name ?? 'unknown'}`
export const hasPoolerIdentity = (pooler: Pick<Multipooler, 'id'>) =>
pooler.id?.cell !== undefined && pooler.id?.name !== undefined
/**
* `routingState.role` is the authoritative writable signal; the deprecated
* `type` field is derived and only used as a fallback when the routing state is
* missing. The role is omitted entirely when it is ROUTING_ROLE_UNKNOWN
* (proto3 zero value).
*/
export const isPrimaryPooler = (pooler: Multipooler) => {
const role = pooler.routingState?.role
if (role !== undefined) return role === 'ROUTING_ROLE_PRIMARY'
return pooler.type === 'PRIMARY'
}
/**
* Health as reported by the pooler's topology record — not a live probe: a
* pooler that crashes without publishing STOPPING/SHUTDOWN can leave an
* ACTIVE/SERVING record behind until the topology evicts it, and "healthy"
* says nothing about replication progress. Live signals (WAL receiver state,
* replay position) exist on `GET /poolers/{cell}/{name}/status` but require a
* per-pooler fanout.
*/
export const getPoolerStatus = (pooler: Multipooler): HaPoolerStatus => {
// Lifecycle values may arrive with or without the proto enum prefix.
const lifecycle = (pooler.lifecycleStatus?.status ?? '').replace(/^LIFECYCLE_/, '')
if (lifecycle === 'QUARANTINED' || lifecycle === 'SHUTDOWN') return 'unhealthy'
if (lifecycle === 'STARTING') return 'coming_up'
if (lifecycle === 'STOPPING') return 'going_down'
// Lifecycle is ACTIVE or unknown: fall back to the serving status. An absent
// servingStatus means SERVING (proto3 zero value), i.e. the node is taking
// traffic, so the default reads as healthy.
if (pooler.servingStatus === 'DRAINING') return 'going_down'
if (pooler.servingStatus === 'DISABLED') return 'unhealthy'
return 'healthy'
}
// Matches the status vocabulary of the read replica surfaces (getStatusLabel).
export const HA_POOLER_STATUS_LABELS: Record<HaPoolerStatus, string> = {
healthy: 'Healthy',
coming_up: 'Coming up',
going_down: 'Going down',
unhealthy: 'Unhealthy',
}
const AWS_AZ_REGEX = /\b[a-z]{2}(?:-[a-z]+)+-\d[a-z]\b/
/**
* Cells map 1:1 to availability zones in the alpha, but the exact cell naming
* format is unconfirmed — extract an AZ-shaped substring when there is one and
* fall back to the raw cell name otherwise.
*/
export const formatCellAsAvailabilityZone = (cell: string | undefined) => {
if (!cell) return undefined
return AWS_AZ_REGEX.exec(cell)?.[0] ?? cell
}
// Routing-rule terms are proto int64s, serialized as strings in JSON and
// omitted when zero. Failover counts stay far below Number's safe range.
const parseTerm = (value: string | undefined) => {
const parsed = Number(value ?? 0)
return Number.isFinite(parsed) ? parsed : 0
}
// Orders two primary claimants by routing rule: (coordinator term, leader
// subterm), greatest wins.
const compareRoutingRules = (a: Multipooler, b: Multipooler) => {
const ruleA = a.routingState?.rule
const ruleB = b.routingState?.rule
return (
parseTerm(ruleA?.coordinatorTerm) - parseTerm(ruleB?.coordinatorTerm) ||
parseTerm(ruleA?.leaderSubterm) - parseTerm(ruleB?.leaderSubterm)
)
}
export const buildHaTopology = ({
gateways,
poolers,
}: {
gateways: Multigateway[]
poolers: Multipooler[]
}): HaTopology => {
// The /ha-admin passthrough can return the same record multiple times (one
// copy per cell it fans out to), so both lists must be deduped by id —
// duplicate poolers would otherwise produce duplicate React Flow node ids,
// and extra copies of the primary would render as replicas. Records without
// a complete identity can't be told apart, so they are never deduped (keying
// by the record itself keeps each one unique).
const uniqueGateways = uniqBy(gateways, (gateway) =>
hasPoolerIdentity(gateway) ? getPoolerKey(gateway) : gateway
)
const uniquePoolers = uniqBy(poolers, (pooler) =>
hasPoolerIdentity(pooler) ? getPoolerKey(pooler) : pooler
)
const sortedPoolers = [...uniquePoolers].sort((a, b) =>
getPoolerKey(a).localeCompare(getPoolerKey(b))
)
const poolersByShard = groupBy(
sortedPoolers,
(pooler) =>
`${pooler.shardKey?.database ?? ''}/${pooler.shardKey?.tableGroup ?? ''}/${pooler.shardKey?.shard ?? ''}`
)
const shards = Object.entries(poolersByShard)
.sort(([a], [b]) => a.localeCompare(b))
.map(([id, shardPoolers], index) => {
const [primaries, replicas] = partition(shardPoolers, isPrimaryPooler)
// During a failover the outgoing and incoming primary can briefly both
// claim ROUTING_ROLE_PRIMARY — the highest routing rule wins, matching
// the multigateway's election. Losing claimants render as replicas
// rather than being dropped; ties keep the first in sorted order so the
// result stays deterministic.
const primary = primaries.reduce<Multipooler | undefined>(
(best, candidate) =>
best === undefined || compareRoutingRules(candidate, best) > 0 ? candidate : best,
undefined
)
return {
id,
name: `Shard ${index + 1}`,
primary,
replicas: [...primaries.filter((pooler) => pooler !== primary), ...replicas],
}
})
return { gateways: uniqueGateways, shards }
}
/**
* Query `select` projecting poolers down to the fields the topology depends on
* (identity, shard, routing role) — live status is self-fetched by the
* individual nodes and edges. React Query's structural sharing then keeps the
* result referentially stable across polls, so refetches only re-run
* layout/fitView when the topology actually changes (volatile fields like
* lifecycle timestamps would otherwise churn the data identity on every poll).
*/
const projectRoutingRule = (rule: NonNullable<Multipooler['routingState']>['rule']) =>
rule === undefined
? undefined
: { coordinatorTerm: rule.coordinatorTerm, leaderSubterm: rule.leaderSubterm }
const projectRoutingState = (routingState: Multipooler['routingState']) =>
routingState === undefined
? undefined
: { role: routingState.role, rule: projectRoutingRule(routingState.rule) }
export const selectTopologyPoolers = (data: HaClusterPoolersData): Multipooler[] =>
(data.poolers ?? []).map((pooler) => ({
id: pooler.id === undefined ? undefined : { cell: pooler.id.cell, name: pooler.id.name },
shardKey:
pooler.shardKey === undefined
? undefined
: {
database: pooler.shardKey.database,
tableGroup: pooler.shardKey.tableGroup,
shard: pooler.shardKey.shard,
},
routingState: projectRoutingState(pooler.routingState),
type: pooler.type,
}))