Files
supabase/apps/www/md-alternates.test.ts
Pamela Chia 9113c2ba04 feat: markdown alternate tags + llms.txt cleanup (#48287)
The June/July marketing redesign (#47271, #47228) rebuilt the homepage
and product pages off the Pages Router, silently dropping their `<link
rel="alternate" type="text/markdown">` head tags, and llms-full.txt has
been accidentally embedding every blog/customer/event page via an
`MD_CONTENT` spread. I restored the tags behind a shared helper, added a
CI drift test so a future redesign can't drop them silently again, and
trimmed both llms files to the agreed docs-index shape.

**Changed:**

- **Markdown siblings advertised again**: homepage, the 5 product pages,
pricing, and blog emit absolute `.md` alternate URLs via a new
`mdAlternates(slug)` helper (the one documented consumer of the tag
parses it from `<head>` and fetches the `.md` sibling, so tags must
point at the sibling, never the page itself).
- **Drift test**: a vitest file walks `content/md/**` and asserts every
markdown-served slug's page wires the helper (or is covered by the Pages
Router `_app.tsx` mechanism, whose alternate-link wiring the test also
asserts directly so removing it fails CI too). Source-level assertions
by design: page modules can't be imported under www's vitest config.
Fails correctly when wiring is removed (verified by hiding a page and by
altering the `_app.tsx` tag).
- **Vector orphan fixed**: `content/md/vector.md` moved to
`modules/vector` matching the live route (the page previously had no
negotiation or tag, and `/modules/vector.md` 404'd); `/vector.md` now
308s to `/modules/vector.md` and the legacy `/llms/vector.txt` redirect
no longer chains.
- **llms.txt + llms-full.txt**: the `## Product Overview` sections are
gone from both, each keeps a `## Pricing` section. This deletes the
hand-maintained links array (a drift trap) and fixes the accidental
~470-page embed, shrinking llms-full.txt from ~9.8MB to ~4.9MB and
dropping the 4.1MB generated content module from that route's serverless
bundle.

**Note:** this PR is scoped to apps/www only. The docs side
(troubleshooting pages and the rest of the docs surface) is handled
separately through a consolidated manifest-gated mechanism; an earlier
troubleshooting-tag commit was reverted out of this branch to keep the
scopes clean.

<details>
<summary>Why alternate tags matter (background)</summary>

Agents ingest markdown far more efficiently than our rendered HTML: a
fraction of the tokens and no extraction step. Since #47770 removed
UA-based serving (UA sniffing broke a major AI app's fetcher and
poisoned CDN caches), markdown is served only on explicit request: a
`.md` suffix URL, an `Accept: text/markdown` header, or llms.txt. That's
the right serving model, but it makes the markdown twin invisible to any
agent that doesn't already know our URL convention, and the major AI
fetchers send browser/wildcard Accept headers, so bare URLs hand them
HTML.

The `<link rel="alternate" type="text/markdown">` head tag is the
standards-based advertisement of the sibling. It has a documented
consumer today: an agent CLI that parses the tag from `<head>` and then
fetches the `.md` sibling, which is also why the tag must point at a
real sibling URL and never at the page itself. Peer docs sites ship this
tag as table stakes. These www pages used to carry it until the
June/July marketing redesign silently dropped it; the drift test in this
PR turns that regression class into a CI failure.

</details>

## To test

Tested locally (www + docs dev servers):
- [x] `/llms.txt` renders `## Documentation` + single-link `## Pricing`,
no Product Overview
- [x] `/llms-full.txt` renders `# Supabase` → `## Pricing` → `##
Documentation`, no Product Overview, ~4.9MB
- [x] Full www suite: 6 files / 71 tests green; drift test fails
correctly when a page is removed or the `_app.tsx` wiring is altered
- [x] `generateMdContent.mjs` emits `modules/vector`, bare `vector` slug
gone

On the Vercel preview (browser-verified with Playwright):
- [x] Alternate tag present on `/`, `/auth`, `/database`, `/storage`,
`/edge-functions`, `/realtime`, `/pricing`, and a blog post: exactly one
tag each, href = preview origin + `.md` sibling
- [x] `/vector.md` → 308 → `/modules/vector.md`, renders as markdown (`#
Supabase Vector`)
- [x] `/llms.txt` shows single-link `## Pricing`, no Product Overview
- [x] Coverage sweep: all 482 `MD_PAGES` slugs + changelog index/entry
curled on the preview; 471 pages carry exactly one tag, all `.md`
siblings 200 as `text/markdown`. The 11 misses are legacy blog slugs
whose HTML 308-redirects away (stale `MD_PAGES` entries predating this
PR, no head to tag; follow-up tracked in Linear)

Post-merge prod:
- [ ] Full llms.txt link sweep (every linked URL 200s; previews can't
cover the docs-hosted links)

## Linear

- fixes GROWTH-1013


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Markdown alternate links across key product, pricing, blog, and
troubleshooting pages.
* Added Supabase Vector documentation covering features, use cases,
workflows, and technical details.
* Updated AI-focused documentation indexes with dedicated pricing
content.
  * Added redirects for updated Vector documentation URLs.

* **Tests**
* Added coverage to verify Markdown documentation links stay aligned
with available pages.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 17:06:16 +08:00

110 lines
4.2 KiB
TypeScript

import { existsSync, promises as fs } from 'node:fs'
import path from 'node:path'
import { describe, expect, it } from 'vitest'
import { SITE_ORIGIN } from './lib/constants'
import { mdAlternates } from './lib/md-alternates'
const DYNAMIC_SLUGS = ['pricing']
const MDX_SECTIONS = ['blog', 'customers', 'events']
async function collectContentMdSlugs(dir: string, prefix = ''): Promise<string[]> {
const dirents = await fs.readdir(dir, { withFileTypes: true })
const slugs: string[] = []
for (const dirent of dirents) {
const slug = prefix ? `${prefix}/${dirent.name}` : dirent.name
if (dirent.isDirectory()) {
slugs.push(...(await collectContentMdSlugs(path.join(dir, dirent.name), slug)))
} else if (dirent.name.endsWith('.md')) {
slugs.push(slug.replace(/\.md$/, ''))
}
}
return slugs
}
async function collectAppRouterPages(
dir: string,
segments: string[] = []
): Promise<Map<string, string>> {
const pages = new Map<string, string>()
const dirents = await fs.readdir(dir, { withFileTypes: true })
for (const dirent of dirents) {
if (dirent.isDirectory()) {
if (dirent.name.startsWith('[') || dirent.name.startsWith('_')) continue
const nextSegments = dirent.name.startsWith('(') ? segments : [...segments, dirent.name]
const nested = await collectAppRouterPages(path.join(dir, dirent.name), nextSegments)
nested.forEach((filePath, slug) => pages.set(slug, filePath))
} else if (dirent.name === 'page.tsx') {
pages.set(segments.join('/') || 'homepage', path.join(dir, dirent.name))
}
}
return pages
}
describe('mdAlternates', () => {
it('emits the absolute .md sibling URL', () => {
expect(mdAlternates('auth')).toEqual({
types: { 'text/markdown': `${SITE_ORIGIN}/auth.md` },
})
expect(mdAlternates('modules/vector')).toEqual({
types: { 'text/markdown': `${SITE_ORIGIN}/modules/vector.md` },
})
})
})
describe('markdown alternate drift', () => {
it('every markdown-served slug has a page advertising its .md sibling', async () => {
const contentSlugs = await collectContentMdSlugs(path.join(process.cwd(), 'content/md'))
const expectedSlugs = [...contentSlugs, ...DYNAMIC_SLUGS]
const appPages = await collectAppRouterPages(path.join(process.cwd(), 'app'))
expect(contentSlugs.length).toBeGreaterThan(0)
for (const slug of expectedSlugs) {
const appPagePath = appPages.get(slug)
if (appPagePath) {
const source = await fs.readFile(appPagePath, 'utf-8')
expect(
source.includes(`alternates: mdAlternates('${slug}')`),
`${path.relative(process.cwd(), appPagePath)} must contain "alternates: mdAlternates('${slug}')" in its metadata export`
).toBe(true)
} else {
expect(
existsSync(path.join(process.cwd(), 'pages', `${slug}.tsx`)),
`no page found for markdown slug "${slug}" — App Router pages need mdAlternates, Pages Router pages are covered by _app.tsx`
).toBe(true)
}
}
})
it('pages/_app.tsx advertises the .md sibling for Pages Router pages', async () => {
const source = await fs.readFile(path.join(process.cwd(), 'pages', '_app.tsx'), 'utf-8')
expect(
source.includes('MD_PAGES.has('),
'pages/_app.tsx must gate the markdown alternate on MD_PAGES membership'
).toBe(true)
expect(
source.includes('rel="alternate" type="text/markdown"'),
'pages/_app.tsx must render the text/markdown alternate link for markdown-served slugs'
).toBe(true)
})
it.for(MDX_SECTIONS)('%s pages advertise their .md sibling', async (urlPrefix) => {
const appPagePath = path.join(process.cwd(), 'app', urlPrefix, '[slug]', 'page.tsx')
if (!existsSync(appPagePath)) {
expect(
existsSync(path.join(process.cwd(), 'pages', urlPrefix)),
`section "${urlPrefix}" has neither app/${urlPrefix}/[slug]/page.tsx nor pages/${urlPrefix}/`
).toBe(true)
return
}
const source = await fs.readFile(appPagePath, 'utf-8')
const wiring = 'alternates: mdAlternates(`' + urlPrefix + '/${slug}`)'
expect(
source.includes(wiring),
`app/${urlPrefix}/[slug]/page.tsx must contain "${wiring}" in its generateMetadata`
).toBe(true)
})
})