From 95d1e8abe8b2c946057d8d84b8824a5cdd14e2fe Mon Sep 17 00:00:00 2001 From: Pamela Chia Date: Mon, 11 May 2026 15:14:36 +0800 Subject: [PATCH] fix(www): drop malformed legal URLs from sitemap_www.xml (#45775) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ## Summary Google Search Console flagged 4 "URL not allowed" errors on `sitemap_www.xml` — malformed URLs like `https://supabase.comdata/legal/terms/v1` (missing slash, non-existent path). The generator was globbing `data/**/*.mdx`, picking up the 4 content-source MDX files under `data/legal/` that are imported by `pages/terms.tsx` and `pages/enterprise-terms.tsx` but are not themselves routed. With no path replacement mapping `data/...` to a route and no leading slash, the URL template concatenated to garbage. The real `/terms` and `/enterprise-terms` URLs come from the `pages/*.tsx` glob and are unaffected. ## Changes - Remove `data/**/*.mdx` glob (and its companion `!data/*.mdx` exclude) from the sitemap generator. `apps/www/data/` has no routed MDX, only content sources imported into pages. - Anchor the `pages` prefix replace: `.replace('pages', '')` → `.replace(/^pages/, '')`. String-form replace is first-occurrence and would mangle any future filename containing `pages` as a non-prefix substring (e.g., `_blog/about-pages.mdx` → `/blog/about-`). No current files trigger this; defensive hardening. ## Testing Regenerated the sitemap locally and verified: - [x] `grep -c "supabase.comdata" public/sitemap_www.xml` → `0` (was 4) - [x] `https://supabase.com/terms` and `https://supabase.com/enterprise-terms` still present - [x] Every `` matches `^https://supabase\.com(/[a-zA-Z0-9].*)?$` (no malformed URLs of any kind) - [x] Total loc count stable across both commits (regression-free for the anchor change) Local count is lower than prod (527 vs 906) because `.next/server/pages/**` partner/expert/feature HTML globs only resolve after a full build — runs correctly via `postbuild` on Vercel. After deploy lands, resubmit `sitemap_www.xml` in Google Search Console to force a re-crawl (otherwise daily-ish). Expect status to flip from "4 errors" to "Success" and Discovered pages: 906 → 902. ## Linear - fixes GROWTH-837 ## Summary by CodeRabbit ## Release Notes * **Chores** * Improved sitemap generation to properly index specific content sections (blog, case studies, customers, events, and alternatives) with refined route path processing for better search engine discoverability. [![Review Change Stack](https://storage.googleapis.com/coderabbit_public_assets/review-stack-in-coderabbit-ui.svg)](https://app.coderabbit.ai/change-stack/supabase/supabase/pull/45775) --- apps/www/internals/generate-sitemap.mjs | 8 ++++---- 1 file changed, 4 insertions(+), 4 deletions(-) diff --git a/apps/www/internals/generate-sitemap.mjs b/apps/www/internals/generate-sitemap.mjs index 8bca05cc71d..8b22ea7cd00 100644 --- a/apps/www/internals/generate-sitemap.mjs +++ b/apps/www/internals/generate-sitemap.mjs @@ -10,13 +10,11 @@ async function generate() { 'pages/*.tsx', 'pages/*.mdx', 'pages/**/*.tsx', - 'data/**/*.mdx', '_blog/*.mdx', '_case-studies/*.mdx', '_customers/*.mdx', '_events/*.mdx', '_alternatives/*.mdx', - '!data/*.mdx', '!pages/_*.js', '!pages/_*.tsx', '!pages/api', @@ -38,7 +36,7 @@ async function generate() { .map((page) => { const path = page .replace('.next/server/pages', '') - .replace('pages', '') + .replace(/^pages/, '') .replace('.html', '') // add a `/` for blog posts .replace('_blog', `/${blogUrl}`) @@ -120,7 +118,9 @@ async function generate() { const changelogDetailUrls = (() => { try { const rss = readFileSync('public/changelog-rss.xml', 'utf-8') - const matches = [...rss.matchAll(/(https:\/\/supabase\.com\/changelog\/\d+[^<]*)<\/link>/g)] + const matches = [ + ...rss.matchAll(/(https:\/\/supabase\.com\/changelog\/\d+[^<]*)<\/link>/g), + ] const uniqueUrls = [...new Set(matches.map((match) => match[1]))] return uniqueUrls.map(