14 Commits
Author SHA1 Message Date
Pamela Chia 21a27eeb4f feat(www): canonicalize homepage markdown at /index.md (#49384)
The www root markdown lived at an accidental URL: `/.md` served the
homepage markdown only because middleware strips the `.md` suffix and
the empty slug fell through to the homepage allowlist entry, while the
canonical-looking `/index.md` 404'd. The served markdown also opened
with stale legacy positioning copy that no longer matches the site. I
renamed the homepage content slug to `index` end-to-end so `/index.md`
is the one canonical markdown URL.

**Changed:**
- **`/index.md` serves the homepage markdown (200 `text/markdown`)**:
`content/md/homepage.md` renamed to `index.md`; the middleware bare-root
slug mapping, the generator's sort special-case, and the homepage
alternate tag follow, so the tag now advertises `/index.md`.
- **Legacy aliases 308 to the canonical URL**: `/.md`, `/homepage.md`,
and bare `/index` redirect via `lib/redirects.js`; `/llms/homepage.txt`
retargeted straight to `/index.md` to avoid a redirect chain. New
`next.config.test.ts` assertions pin all four.
- **Positioning refreshed**: the markdown now opens with "Supabase is
the Postgres development platform" (matching the site title), replacing
the outdated tagline.
- **Generator safety**: the redirect-exclusion filter in
`generateMdContent.mjs` now exempts the `index` slug (its HTML page is
`/`, not `/index`, so a `/index` redirect never refers to it), and the
build fails if `content/md/index.md` ever goes missing while middleware
still maps `/` to the `index` slug.
- **CI actually runs the new assertions**: I widened the `www-tests.yml`
paths filter to include `apps/www/lib/**/*.js`,
`apps/www/content/md/**`, and `apps/www/scripts/**/*.mjs`. It previously
only matched `.ts*` and the next.config files, so a PR touching only
`lib/redirects.js`, the markdown content, or the generator would skip
the tests that pin these redirects.

**Note:** the existing homepage alternate tag still exists, re-pointed
to the canonical URL. Whether the homepage should advertise a markdown
sibling at all is a separate decision; leaving it aimed at a 308 would
break tag consumers. Positioning wording is editorial, happy to tweak.

## To test
Tested on Vercel preview:
- [x] `curl -si <preview>/index.md`: expect 200 `content-type:
text/markdown`, body opens with the Postgres development platform
positioning and no longer contains the old tagline
- [x] `curl -sI <preview>/.md`: expect 308 with `location: /index.md`
- [x] `curl -sI <preview>/homepage.md` and `curl -sI
<preview>/llms/homepage.txt`: expect 308 with `location: /index.md`
- [x] `curl -sI <preview>/index`: expect 308 with `location: /`
- [x] `curl -s -H "Accept: text/markdown" -o /dev/null -w "%{http_code}
%{content_type}" <preview>/`: expect `200 text/markdown` (bare-URL
negotiation unchanged)
- [x] `curl -s <preview>/ | grep -o 'type="text/markdown"
href="[^"]*"'`: expect href ending `/index.md`

## Linear
- fixes GROWTH-1117



<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

- **New Features**
- Added support for `/index.md` as the canonical Markdown representation
of the homepage.
- Added permanent redirects for legacy homepage Markdown and text URLs.
  - Added `/index` to `/` redirect handling.

- **Bug Fixes**
- Updated homepage metadata, alternate links, Markdown negotiation, and
content generation to consistently use the new canonical path.
  - Improved homepage content description.

- **Tests**
- Expanded coverage for homepage Markdown routes, redirects, and URL
matching.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-08-24 15:19:56 +08:00
Pamela Chia f208743432 fix(www): changelog md negotiation slug-set gate (#49357)
Bare-URL `Accept: text/markdown` negotiation never fires on changelog
entries authored after the GitHub-discussions backfill: the middleware
gate `/^changelog\/\d+/` only matches legacy numeric slugs (from
`legacy_gh_discussion` frontmatter), so agents that signal markdown via
Accept get HTML on every new entry. I found this in the independent
review round on #48475; pre-existing, not introduced there.

**Changed:**
- **Non-legacy entries negotiate markdown**: `generateMdContent.mjs` now
lists `public/changelog/*.md` (written moments earlier by
`generateStaticContent.mjs` in the same `content:build:core` chain) and
emits a `CHANGELOG_PAGES` set into the generated module; the middleware
regex becomes a set lookup, so negotiation coverage derives from the
exact static files served and can't drift from what's published.
- **Unknown and deep changelog paths stop negotiating**: the old regex
prefix-matched paths like `changelog/100/bar` and nonexistent numeric
slugs, rewriting them to missing `.md` files (404 under a markdown
Accept); they now pass through to the dynamic route's canonicalizing
308/404.
- **Build guard**: zero collected changelog slugs on Vercel fails the
build (today a zero-entry changelog fetch ships empty output with a
green build), and a shape assertion fails the build if collected slugs
ever lose the `changelog/` prefix the middleware matches on. Locally
without `CHANGELOG_SYNC_APP_*` secrets it warns and changelog
negotiation is off, matching the absent content.
- **`/changelog` index gated the same way**: the index slug is emitted
into the set only when `public/changelog.md` was generated, replacing
the hardcoded `slug === 'changelog'` branch; locally without secrets the
index no longer rewrites to a nonexistent file.

**Note:** script order in `content:build:core` is load-bearing (static
content generation must precede md content generation); the Vercel guard
turns a reorder into a loud build failure instead of a silent empty
gate.

## To test
Tested on the Vercel preview (`zone-www-dot-com` deployment of head
`f451da3`):
- [x] `curl -sI -H "Accept: text/markdown" <preview>/changelog` and
`curl -sI <preview>/changelog.md`: got 200 `text/markdown` (index via
the generated gate)
- [x] `curl -sI -H "Accept: text/markdown"
<preview>/changelog/pipelines`: got 200 `text/markdown` (prod today
returns `text/html`)
- [x] Same curl against the legacy numeric slug
`48235-migration-of-...`: got 200 `text/markdown` (no regression)
- [x] `curl -sI -H "Accept: application/json"
<preview>/changelog/pipelines`: got 406 (prod today returns 200 HTML)
- [x] Explicit `.md` fetches for both slug shapes
(`/changelog/pipelines.md`, `/changelog/48235-....md`): got 200
`text/markdown`
- [x] `curl -sI -H "Accept: text/markdown"
<preview>/changelog/does-not-exist-xyz`: got a 404 HTML passthrough from
the dynamic route, not a 406
- [x] `pnpm test middleware.test.ts` in `apps/www` at head: 41/41 pass
(36 pre-existing + 5 new). No CI job runs the www vitest suite, so this
local run is the only oracle for the new tests.

## Linear
- fixes GROWTH-1062


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Improved changelog page handling, including markdown versions of
published entries.
- Added content negotiation for supported changelog formats, with clear
responses for unsupported requests.
- **Bug Fixes**
- Prevented unpublished numeric-prefix pages from being treated as
published.
  - Fixed deep links under published changelog entries.
- **Reliability**
- Changelog availability is now detected automatically, with improved
validation during content generation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-08-21 15:01:50 +08:00
Pamela Chia 770f1c2b06 fix(aeo): remove ua-based markdown serving (#47770)
## Summary
The `ChatGPT-User` live-fetch agent's user-facing reader hard-fails
(`(400) OK`) on pages we serve it as markdown via user-agent matching,
which made supabase.com blog and product pages unreadable in that
assistant. I root-caused this with a controlled fetch diagnostic
cross-checked against our request logs: the failing fetches never reach
our origin (the failure is cached on their side), pages served as plain
HTML read fine everywhere we tested, and the same failure reproduces on
other major sites that serve UA-matched markdown, so the reader bug is
upstream.

This PR removes user-agent-based markdown serving entirely rather than
special-casing one agent: UA sniffing is a guess about contractless
clients whose fetchers change without notice, and this incident showed
the failure mode is silent (we keep serving 200s while the user-facing
agent breaks). Markdown remains available on every explicit signal —
`Accept: text/markdown` q-value negotiation, explicit `.md` URLs, and
llms.txt — which is the same contract-driven model the Claude fetcher
already uses successfully (it sends `Accept: text/markdown, text/html,
*/*` and keeps receiving markdown after this change).

## Changes
- Remove the `LLM_USER_AGENT` regex and the `userAgent` parameter from
`negotiateMarkdown` in `packages/common/markdown-negotiation.ts`;
decisions now depend only on `Accept`, the `.md` suffix, and the
markdown-variant manifest
- Update both consuming middlewares (`apps/www`, `apps/docs`) to the new
signature; no behavior change for Accept-negotiated or `.md` requests
- Add the missing `Vary: Accept` header to docs guides-md 200 responses
(the www `api-v2/md` route already declares it)
- Fix a pre-existing www bug surfaced in review: explicit changelog
`.md` URLs rewrote to a doubled `.md.md` path (404) under a
markdown-preferring `Accept`, and 406'd on a non-matching `Accept`. The
www middleware now strips the `.md` suffix before slug lookup and passes
`isMarkdownSuffix` into `negotiateMarkdown`, folding the separate
`MD_PAGES` `.md` block into the single negotiation path (same shape as
the docs middleware)
- Rework tests: UA-independence suites replace the per-agent rewrite
tests; a probe Accept header now 406s regardless of user agent
(previously agent UAs were exempt); new changelog `.md` negotiation
coverage

## Testing
Tested locally:
- [x] www middleware suite 36/36, docs middleware suite 17/17
- [x] typecheck green for common, www, docs

Verified on the Vercel previews (www + docs) with curl:
- [x] `ChatGPT-User` and `Claude-User` UA GETs on blog/pricing/guide
pages return `text/html` with a default Accept
- [x] Claude's real Accept (`text/markdown, text/html, */*`) still
returns `text/markdown`; `Accept: text/markdown` and `.md` URLs return
`text/markdown`; probe Accept returns 406
- [x] `/changelog/<slug>.md` with `Accept: text/markdown` returns the
entry markdown as a direct 200 (production today detours through a 308
to the bare URL); changelog index `.md` and bare-entry Accept
negotiation also verified
- [x] docs guides markdown 200s carry `Vary: Accept`

The intermediate commit (ChatGPT-User-only exclusion) was already
verified on the preview: `ChatGPT-User` got HTML while
`Accept`/`.md`/other-UA markdown was unaffected.

Expected effects post-merge: UA-driven markdown volume in the request
logs (~92% of md traffic) collapses to the Accept + `.md` baseline;
named-agent page requests return to prerendered/static serving,
reversing the extra Vercel function invocations the UA rewrite
introduced; user-facing readability in the affected assistant recovers
within ~24h as its fetch cache revalidates. The md-share dashboard gets
a dated annotation; the ratio is not comparable across this change.

## Linear
- fixes GROWTH-973


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Markdown and HTML routing now depends on the request’s `Accept` header
and `.md` links, making content negotiation more predictable.
* Requests that don’t accept available content now consistently return
`406 Not Acceptable`, even for bot-like user agents.
* Guide markdown responses now include an `Accept`-based cache variation
header to improve correct caching behavior.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-10 13:50:34 +08:00
Pamela Chia cd52669f1f fix(docs): negotiate /guides/* markdown via shared helper (#45432)
## Summary

This brings docs `/guides/*` to full content negotiation for AI agents
(GROWTH-811):
RFC 9110 q-value parsing instead of a `.includes('text/markdown')`
substring match,
a 406 when the client rejects every type the route can produce, and
markdown rewrites
for known LLM user agents.

I implemented it by extracting the negotiation into a shared
`common/markdown-negotiation`
module consumed by both `apps/docs/middleware.ts` and
`apps/www/middleware.ts`, rather than
duplicating the helpers into docs and keeping them in sync by hand with
www (#45394). Single
source of truth, no re-sync burden. www is refactored onto the shared
helper with no behavior
change.

## Changes

### docs `/guides/*` content negotiation (GROWTH-811)

- Replace the `.includes('text/markdown')` substring match with RFC 9110
q-value parsing.
- Return 406 (`Cache-Control: no-store`, `Vary: Accept`) when Accept
excludes every type the
route serves. Bypassed for LLM user agents, the `.md` suffix, and
clients sending no Accept.
- Rewrite to `/api/guides-md/<slug>` for LLM user agents (Claude-User,
Claude-Web, ChatGPT-User,
  PerplexityBot) regardless of Accept.
- Preserve the existing `.md` suffix routing and the entire
`/reference/*` block.

### Shared negotiation helper

- New `packages/common/markdown-negotiation.ts`:
`negotiateMarkdown(signals, route)` returns
`'markdown' | 'not-acceptable' | 'pass'`. Internalizes q-value parsing,
the LLM user-agent
  match, the UA-length cap, and the markdown-vs-html preference.
- `apps/www/middleware.ts`: refactored to consume the shared helper; its
duplicated copy of the
negotiation helpers (added in #45394) is removed. `.md` early-return,
changelog routing, and
first-referrer cookie stamping are unchanged (no behavior change,
covered by its existing tests).

### Tests

- New `apps/docs/middleware.test.ts`: q-value priority, the 406 path,
`.md` suffix, LLM UA
override, browser default Accept, training-crawler and substring-embed
exclusion, and the
  `/reference/*` exemption.
- New `packages/common/markdown-negotiation.test.ts`: the same decision
matrix at the unit level
(q-values, 406, LLM UAs, `.md`, `*/*`, training crawlers, OWS,
out-of-range q).

## Testing (Vercel preview)

After Vercel posts a preview URL, save it once then run the probe set.

```bash
echo 'PREVIEW_HOST' > /tmp/growth-811-host.txt
HOST=$(cat /tmp/growth-811-host.txt)

# 1) Browser-style Accept -> HTML 200
curl -sI -A "Mozilla/5.0" \
  -H 'Accept: text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8' \
  "https://$HOST/docs/guides/auth"

# 2) Accept: text/markdown -> markdown 200
curl -sI -H 'Accept: text/markdown' "https://$HOST/docs/guides/auth"

# 3) text/html;q=1.0, text/markdown;q=0.5 -> HTML 200
curl -sI -H 'Accept: text/html;q=1.0, text/markdown;q=0.5' "https://$HOST/docs/guides/auth"

# 4) unsupported Accept -> 406 + Cache-Control: no-store + Vary: Accept
curl -sI -H 'Accept: application/x-content-negotiation-probe' "https://$HOST/docs/guides/auth"

# 5) User-Agent: Claude-User/1.0 (any Accept) -> markdown 200
curl -sI -A 'Claude-User/1.0' "https://$HOST/docs/guides/auth"
```

### After merge

Run
[acceptmarkdown.com/readiness-check](https://acceptmarkdown.com/readiness-check)
against `https://supabase.com/docs/guides/auth`: expect 100/100.

## Linear

- fixes GROWTH-811
2026-06-09 17:49:35 +08:00
Pamela Chia dff4744805 fix(www): respect Accept q-values and 406 unsupported types (#45394) 2026-05-01 14:47:33 +09:00
Francesco Sansalvadore 8ba1054dfe chore(www): changelog formatting (#45364)
- change changelog.md formatting
- make changelog entries slugs more descriptive (eg
/changelog/123-new-change)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Refactor**
* Updated changelog entry URLs to use slug-based identifiers instead of
numeric IDs for improved readability and SEO-friendliness, with
automatic redirects for existing links.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-04-29 13:56:32 +02:00
Francesco SansalvadoreandCopilot Autofix powered by AI 580598f0e8 feat(www): update changelog layout, rss and md files (#45219)
- Update Changelog [index page
layout](https://zone-www-dot-com-git-feat-changelog-update-supabase.vercel.app/changelog):
  - with full timeline
  - filterable based on text search and tags
- New Changelog [detail
pages](https://zone-www-dot-com-git-feat-changelog-update-supabase.vercel.app/changelog/45071)
  - all added to www_sitemap
- Changelog [RSS
Feed](https://zone-www-dot-com-git-feat-changelog-update-supabase.vercel.app/changelog/45071)
+ llm-friendly
[/changelog.md](https://zone-www-dot-com-git-feat-changelog-update-supabase.vercel.app/changelog.md)
- and llm-friendly changelog detail md files:
https://zone-www-dot-com-git-feat-changelog-update-supabase.vercel.app/changelog/45071.md

## Before
<img width="1604" height="1094" alt="Screenshot 2026-04-27 at 17 07 55"
src="https://github.com/user-attachments/assets/eac52f14-e447-4f64-8d50-a8e287ccf989"
/>

## After
<img width="1247" height="849" alt="changelog-index"
src="https://github.com/user-attachments/assets/69b7bae1-63eb-4a4d-a065-7541ed9738b4"
/>

### Detail page
<img width="1695" height="1101" alt="Screenshot 2026-04-27 at 18 27 27"
src="https://github.com/user-attachments/assets/accd4be8-d665-43ed-bcb7-0e6baf537762"
/>


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

## Release Notes

* **New Features**
* Redesigned changelog page with full-text search and product tag
filtering
  * Individual pages for each changelog entry with dedicated URLs
  * Added RSS feeds for changelog updates and product-specific feeds
  * Copy changelog entries as markdown with one click
  * Direct sharing integration with ChatGPT and Claude

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Copilot Autofix powered by AI <62310815+github-advanced-security[bot]@users.noreply.github.com>
2026-04-29 12:31:30 +02:00
Pamela Chia a98a4928b4 feat(www): rewrite to md for known llm user agents (#45328) 2026-04-29 02:41:24 +09:00
Pamela Chia d409836ca7 feat(www,docs): serve marketing pages as .md, advertise via link rel=alternate (#45277)
## Summary

Adds `/<page>.md` routes for 10 marketing/product pages (homepage, auth,
database, edge-functions, realtime, storage, vector, pricing,
modules/cron, modules/queues) so AI agents can fetch clean markdown
instead of parsing JS-rendered HTML. Also advertises the markdown
alternate via `<link rel="alternate" type="text/markdown">` on marketing
and docs pages so agents can discover it.

Pricing is generated dynamically via `generatePricingContent()` (single
source of truth with `/llms.txt` and `/llms-full.txt`); the other nine
slugs are bundled at build time from `content/md/*.md` into a
`MD_CONTENT` map.

Supersedes #44891 (rebased fresh off current master to avoid a 9-commit
replay over rename/rename conflicts created by #44897).

## Changes

- New `/api-v2/md/[...slug]` route handler returns the bundled markdown
(or dynamic pricing) with `Content-Type: text/markdown`,
`X-Content-Type-Options: nosniff`, and appropriate cache headers
- Middleware rewrites `/<slug>.md` and `Accept: text/markdown` to the
API route for the `MD_PAGES` allowlist; trailing-slash variants
(`/auth/`) are normalized so they resolve the same as `/auth`
- Build-time codegen `scripts/generateMdContent.mjs` scans `content/md/`
and emits `app/api-v2/md/content.generated.ts` exporting both
`MD_CONTENT` (Map) and `MD_PAGES` (Set, incl. dynamic `pricing`). Fails
the build on slug collision between `content/md/` and `DYNAMIC_SLUGS`.
Adding a new marketing `.md` is just dropping a file in `content/md/`
(also update `PRODUCT_OVERVIEW_LINKS` in `/llms.txt` since that list is
editorial).
- 8 permanent redirects `/llms/<product>.txt` → `/<product>.md` so
legacy URLs in caches and downstream `llms.txt` copies keep working
- `/llms.txt` product overview now references `.md` URLs (incl.
`modules/cron`, `modules/queues`); `/llms-full.txt` iterates
`MD_CONTENT.values()` (homepage first, then alphabetical) and appends
dynamic pricing
- `/llms/[slug]` route slimmed to proxy SDK reference files (`js.txt`,
`dart.txt`, etc.) since redirects handle product slugs and pricing;
pricing branch retained as fallback in case redirects are bypassed
- `apps/www/pages/_app.tsx` injects the alternate link conditionally
based on `MD_PAGES`; `/pricing` (app router) sets it via page metadata
- `apps/docs/app/page.tsx` (the `/docs` root) sets the text/markdown
alternate to `/llms-full.txt`; per-guide pages override with their
specific `.md` URL via `genGuideMeta` in `GuidesMdx.utils.tsx`. Other
docs pages (reference, troubleshooting) inherit nothing.
- `apps/www/.vercelignore`: replaces the prior `*.md`/`README.md` rules
with `*.md` + `!content/md/**/*.md` so Edge Function READMEs and future
scratch `.md` files aren't silently shipped to the build artifact
- Drops `apps/www/data/llms/*.txt` and the related
`outputFileTracingIncludes`
- Test coverage for the new middleware branches: `.md` suffix rewrite
(allowlisted vs. fall-through), `Accept: text/markdown` content
negotiation, trailing-slash normalization

## Testing (Vercel preview)

Local dev server smoke tests passing on `:3771` after each iteration.
Re-verified on the preview URL after the latest hardening commit:

- [x] `curl -I https://<preview>/llms/auth.txt` — expect `308 Permanent
Redirect` to `/auth.md`
- [x] `curl https://<preview>/auth.md | head -3` — expect `# Supabase
Auth`
- [x] `curl https://<preview>/pricing.md | head -3` — expect `# Supabase
Pricing` with current tier values
- [x] `curl https://<preview>/modules/cron.md | head -3` — expect `#
Supabase Cron`
- [x] `curl -H 'Accept: text/markdown' https://<preview>/ | head -3` —
expect `# Supabase` (homepage.md)
- [x] `curl https://<preview>/llms.txt` — Product Overview section lists
`.md` URLs and includes Cron + Queues
- [x] `curl https://<preview>/llms-full.txt | grep -E '^# Supabase
(Cron\|Queues\|Pricing)'` — Cron and Pricing each match once; Queues
matches twice (marketing module + existing docs guide)
- [x] View source on `/`, `/pricing`, `/database` — expect `<link
rel="alternate" type="text/markdown" href="/<slug>.md">`
- [x] View source on `/docs` — expect `<link rel="alternate"
type="text/markdown" href="/llms-full.txt">`
- [x] View source on a docs guide page (e.g., `/docs/guides/auth`) —
expect per-guide `.md` alternate; reference/troubleshooting pages should
NOT emit a markdown alternate
- [x] `curl -I https://<preview>/auth.md` — expect
`X-Content-Type-Options: nosniff`
- [x] `curl -I -L -H 'Accept: text/markdown' https://<preview>/auth/` —
should resolve to markdown content (trailing-slash normalization, with
Vercel's auto-redirect)

## Linear

- fixes GROWTH-760

## Follow-up (separate PR)

GROWTH-760 also asks about extending `.md` to blog/customers/events.
Different mechanism (path-prefix middleware, MDX read at request time
via `gray-matter`) so it deserves its own review. Will open a follow-up
PR after this lands.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Serve prebuilt and dynamic Markdown docs via new markdown endpoints
and routing; pages now advertise markdown alternates (including
pricing).
  * Added Cron and Queues module documentation pages.

* **Documentation**
  * Minor formatting tweaks to Realtime and Storage docs.

* **Chores**
* Added build-time Markdown content generation and adjusted
ignore/deploy rules for generated files.
* Added redirects from legacy text-based product URLs to new markdown
pages.

* **Tests**
* Expanded tests for markdown routing and content-negotiation behavior.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-04-28 16:41:03 +09:00
Sean Oliver b041ebfaa5 feat(www): Phase 2 — enable cookie stamping on /dashboard and /docs paths (#43677) 2026-03-12 15:23:07 -07:00
Sean Oliver 8ebbad3a5b feat(growth): expand www middleware to /dashboard and /docs (Phase 1 - instrumentation only) (#43413)
## Problem

The `_sb_first_referrer` cookie isn't working. The www middleware
matcher explicitly excludes `/dashboard` and `/docs`, so the cookie
never gets stamped for Studio or Docs traffic. PostHog confirmed: only 1
event with `first_referrer_cookie_present=true` out of ~46.5M Studio
pageviews in the last 7 days.

## Background: what the matcher does

In Next.js, the `matcher` config controls which incoming requests the
middleware function even runs on. If a path doesn't match, the
middleware is skipped entirely — the request passes through untouched.
If it matches, the middleware runs and can mutate the response (set
cookies, headers, etc.).

This matters because Studio's SPA navigation works via silent
`/_next/data/` JSON fetches. If middleware runs on those requests and
returns `NextResponse.next()` with any mutations, it breaks those
fetches and causes full page reloads instead of client-side transitions.

## What we tried before

| PR | www runs on `/dashboard`? | Studio `proxy.ts` runs on all routes?
| Result |
|---|---|---|---|
| **#42768** (Attempt 1) | ✅ Yes — and also intercepts `_next/data` | ✅
Yes — `matcher` config removed, stamps cookie everywhere | Full page
reloads in Studio |
| **#43129** (Full revert) | ❌ No — www middleware deleted entirely | ❌
No — restored to `matcher: '/api/*'` only | Back to baseline, no cookie
stamping anywhere |
| **#43153** (Attempt 2) | ❌ No — `/dashboard` explicitly excluded | ✅
Yes — `matcher` config removed again, stamps cookie everywhere | Full
page reloads in Studio again |
| **#43189** (Attempt 3) | ❌ No — same as #43153 | ✅ Yes — `matcher`
config still removed, cookie stamping made conditional | Still broken |
| **#43190** (Ivan's fix) | ❌ No — `/dashboard` still excluded | ❌ No —
restored to `matcher: '/api/*'` only | Works — but cookie never stamps
for `/dashboard` traffic |
| **#43413** (this PR) | ✅ Yes — sets diagnostic cookie only, no
attribution stamping yet | ❌ No — unchanged, still `matcher: '/api/*'`
only | ❓ Untested in prod |

The common factor in every failure: Studio's `proxy.ts` ran on all
routes (including `_next/data` requests), which broke SPA navigation.
This PR is the first one that runs www middleware on `/dashboard` while
Studio's `proxy.ts` stays in its original narrow `/api/*` scope.

## What changed

This is Phase 1 of a two-phase rollout. We remove `dashboard|docs` from
the matcher's negative lookahead so www middleware runs on those paths —
but instead of stamping cookies, we set a short-lived (60s) diagnostic
cookie `_sb_mw_diag` on `/dashboard` and `/docs` requests.

The diagnostic cookie encodes
`hit=1&would_stamp={0|1}&has_cookie={0|1}`, which Studio telemetry reads
on the initial pageview and reports to PostHog as `mw_diag_hit`,
`mw_diag_would_stamp`, and `mw_diag_has_existing_cookie` properties.

A cookie rather than a header because response headers aren't readable
by JS. It also tests the actual Set-Cookie mutation path that Phase 2
will use (which is what Next.js issue #41885 is specifically about).

Phase 1 answers two key questions before we commit to Phase 2:
1. Does expanding the matcher break Studio SPA navigation?
2. What % of /dashboard arrivals would get a first-referrer cookie
stamped in Phase 2?

## Phase 2 readiness criteria

**Important caveat**: `mw_diag_*` data reflects consented users only and
may under-represent first-visit anonymous traffic. The Phase 2 decision
should account for this — the actual middleware execution rate is likely
higher than what PostHog reports.

### PostHog query spec

**Middleware execution rate**: Of all Studio initial pageviews on
`/dashboard` or `/docs` paths, what percentage have `mw_diag_hit =
true`? Expected: >= 90%. Below 70% warrants investigation (could
indicate edge caching bypassing middleware, or a matcher configuration
issue).

```
Filter: event = "$pageview" AND (current_url contains "/dashboard" OR current_url contains "/docs")
Breakdown: mw_diag_hit (true vs null/missing)
Metric:    count(mw_diag_hit = true) / count(all) * 100
```

**Would-stamp rate**: Of events with `mw_diag_hit = true`, what
percentage have `mw_diag_would_stamp = true`? This tells us what
percentage of Phase 2 traffic would actually get a cookie stamped. No
hard threshold — unexpected values (< 5% or > 95%) suggest a logic bug
worth investigating before Phase 2.

```
Filter: event = "$pageview" AND mw_diag_hit = true
Breakdown: mw_diag_would_stamp (true vs false)
Metric:    count(mw_diag_would_stamp = true) / count(all) * 100
```

**Existing cookie rate**: Of events with `mw_diag_hit = true`, what
percentage have `mw_diag_has_existing_cookie = true`? This tells us how
many users already have the cookie from a prior www visit.

```
Filter: event = "$pageview" AND mw_diag_hit = true
Breakdown: mw_diag_has_existing_cookie (true vs false)
Metric:    count(mw_diag_has_existing_cookie = true) / count(all) * 100
```

### Go / no-go threshold table

| Signal | Go | Investigate | No-Go |
|---|---|---|---|
| `mw_diag_hit` rate (% of /dashboard+/docs pageviews) | >= 90% | 70-90%
| < 70% |
| SPA navigation errors (Sentry / Vercel logs) | No increase | < 0.1%
increase | > 0.5% increase |
| Middleware p99 latency (Vercel function logs) | < 50ms added |
50-100ms | > 100ms |
| Sample volume in first 24h | > 1,000 events | 100-1,000 (extend
window) | < 100 (insufficient data) |

## Changes

- `apps/www/middleware.ts`: Removed `dashboard|docs` from matcher; added
`isDashboardOrDocs` guard that sets `_sb_mw_diag` diagnostic cookie
instead of stamping attribution
- `apps/www/middleware.test.ts`: 12 tests covering cookie stamping on
www paths, diagnostic cookie encoding for all scenarios (external
referrer, direct nav, internal referrer, existing cookie)
- `packages/common/first-referrer-cookie.ts`: Exported
`MW_DIAG_COOKIE_NAME`, `MwDiagData` type, and `parseMwDiagCookie()`
helper
- `packages/common/telemetry.tsx`: Reads `_sb_mw_diag` on initial Studio
pageview; reports `mw_diag_*` properties to PostHog

## Testing

Unit tests pass (13/13 www, 31/31 first-referrer-cookie). Production
validation needed:

- [ ] Studio SPA navigation works (tab changes, SQL editor, no full page
reloads)
- [ ] PostHog shows `mw_diag_hit = true` on Studio initial pageviews
- [ ] `mw_diag_would_stamp` distribution looks reasonable before
enabling Phase 2
- [ ] Monitor 24h before Phase 2

Ref: GROWTH-625 / GROWTH-668
2026-03-09 11:47:20 -07:00
Sean Oliver 75ec7c6e6b feat(growth): re-land first-referrer cookie attribution with fixed middleware matchers (#43153)
## Summary

Re-lands the first-referrer cookie feature from #42768 (reverted in
#43129) with middleware matcher fixes that prevent Studio traffic
interference.

**Tracks:** [GROWTH-651](https://linear.app/supabase/issue/GROWTH-651)

## What changed

New shared module in `packages/common/first-referrer-cookie.ts` that
handles stamping and parsing a first-referrer cookie (referrer, UTMs,
click IDs, landing URL). Each app's middleware calls
`stampFirstReferrerCookie` on the edge response — www and docs are the
primary entry points, Studio is a fallback for direct visits with UTMs.

On the telemetry side, `handlePageTelemetry` now takes an options object
instead of positional args, reads the cookie on initial pageview, and
overrides the referrer if the cookie captured an external source but the
current referrer is internal (i.e., the user navigated cross-app). Also
sends `first_referrer_cookie_present`/`consumed` properties so we can
observe the handoff in PostHog.

The docs middleware matcher was broadened from `/reference/:path*` to
all docs pages so we stamp cookies site-wide, not just on reference
paths.

## Root cause of original revert

Two layers:

1. **Matcher gap**: www middleware ran on `/dashboard/*` traffic in prod
due to Vercel Multi-Zone architecture (www is the gateway for
`supabase.com`, proxying `/dashboard` → Studio, `/docs` → Docs).
Middleware runs *before* rewrites, so www middleware executed on all
proxied traffic.

2. **`_next/data` interception**: The matcher didn't exclude
`_next/data` paths. Client-side navigation in Next.js fetches JSON via
`/_next/data/...` — middleware intercepted these, returned
`NextResponse.next()` with cookie mutations (which processes through the
middleware response pipeline), and this interfered with the JSON
responses, causing full page reloads in the SQL editor.

## How this PR fixes it

| Fix | Detail |
|---|---|
| Exclude `_next/data` | All three matchers (`www`, `docs`, `studio`)
exclude `_next/data` via negative lookahead |
| Exclude `dashboard` + `docs` from www | www middleware no longer runs
on proxied app traffic |
| `/api/` path guard in Studio | Broadened matcher requires explicit
path check for API route filtering |
| `NextResponse.next()` semantics | Cookie stamping only happens on
matched paths; unmatched paths never enter middleware |

### `NextResponse.next()` vs `undefined` nuance

Returning an explicit `NextResponse.next()` with cookie mutations
processes through Next.js's middleware response pipeline (headers are
merged, cookies are set). Returning `undefined` (i.e. the request never
matches the matcher) lets Next.js handle the request completely
untouched. The matcher exclusions ensure `_next/data` and proxied app
paths never enter middleware at all.

## Testing

- ✅ 22 unit tests for shared cookie utilities (all pass)
- ✅ Studio prod build succeeds, middleware recognized as `ƒ Proxy
(Middleware)`
- ✅ Playwright validation: client-side navigation works across 3 page
transitions, `_next/data` requests return 200 OK without middleware
interception, no full-page reloads
- ❌ www/docs SSG builds require platform backend services (expected —
same as master)
2026-02-25 09:24:32 -08:00
Ivan Vasilov 85e6b1143f chore: Revert "fix: persist first referrer across app boundaries (#42768)" (#43129)
This reverts commit
https://github.com/supabase/supabase/commit/04e63dfb2e00c43f8c57e538b8056aa647f4e5b4
since it was causing the Studio app to rerender the full page on every
link navigation.
2026-02-24 11:56:44 +00:00
04e63dfb2e fix: persist first referrer across app boundaries (#42768)
## Summary

Fixes GROWTH-625.

Preserves first-touch attribution across app boundaries by persisting
external referrer context at the edge and consuming it on Studio's
initial pageview.

When users come from an external source to www/docs and then navigate to
Studio, Studio often only sees the internal `supabase.com` hop. This
change preserves the original external context so first-touch
attribution is retained.

## What changed

- **Shared first-referrer cookie utilities**
(`packages/common/first-referrer-cookie.ts`):
- `isExternalReferrer`, `buildFirstReferrerData`,
`serializeFirstReferrerCookie`, `parseFirstReferrerCookie`
- `hasPaidSignals` — detects click IDs (gclid, fbclid, etc.) and paid
utm_medium values
- `shouldRefreshCookie` — centralizes stamp-or-skip decision for all
apps
- `stampFirstReferrerCookie` — shared middleware helper used by all apps
(extracted from duplicated inline logic)
- **Edge middleware on all apps** — stamps cookie for external visitors,
refreshes on paid signals:
  - `apps/www/middleware.ts` (simplified to use shared helper)
  - `apps/docs/middleware.ts` (simplified to use shared helper)
  - `apps/studio/proxy.ts` (integrated into existing proxy file)
- **Docs middleware matcher** — broadened from `/reference/:path*` to
all non-static paths so the first-referrer cookie is stamped on all docs
pages, not just reference paths
- **Telemetry** — Studio consumes cookie on initial pageview
(`packages/common/telemetry.tsx`). `handlePageTelemetry` refactored from
7 positional params to an options object for readability.
- **Tests** — 22 unit tests covering all utilities and edge cases
(including direct-navigation scenario)

## Behavior

- Writes `_sb_first_referrer` cookie when:
  - cookie is not already set and request has an external referrer, OR
- cookie exists but incoming URL has paid traffic signals (click IDs or
paid utm_medium)
- Cookie: 365-day TTL, `domain=supabase.com`, `sameSite=lax`,
`secure=true` in production
- On first Studio pageview, if current referrer is internal and cookie
has external context:
  - use persisted external referrer
  - apply persisted UTM/click-id/landing-url attribution props
- Measurement properties: `first_referrer_cookie_present`,
`first_referrer_cookie_consumed`

## Manual testing

1. Visit `supabase.com/pricing?utm_source=google&utm_medium=cpc` from an
external referrer (or use DevTools to set a `Referer` header)
2. Check `_sb_first_referrer` cookie is set in Application > Cookies
3. Navigate to Studio (`supabase.com/dashboard`)
4. In PostHog (or browser network tab), verify the first `$pageview`
event has:
   - `first_referrer_cookie_present: true`
   - `first_referrer_cookie_consumed: true`
   - `$utm_source: "google"`, `$utm_medium: "cpc"`
   - `$referrer` points to the external source, not `supabase.com`
5. Verify subsequent route changes do NOT include
`first_referrer_cookie_*` properties

## Review feedback addressed

- Added `secure: true` flag on production cookies (Pam's first comment)
- Fixed inaccurate JSDoc on `utms` field — keys retain `utm_` prefix
(Pam's fourth comment)
- Added test coverage for edge cases: malformed URLs, multi-cookie
headers, http:// referrers (Pam's sixth comment)
- Docs matcher broadening: fast-path exit on cookie-exists check keeps
overhead minimal, exclusion list is correct
- Extracted shared middleware helper to eliminate duplication across 3
apps
- Refactored `handlePageTelemetry` from positional params to options
object
- Removed redundant null check in `hasPaidSignals`
- Added direct-navigation test case
- Deleted dead `apps/learn/middleware.ts`
- Fixed studio build: integrated cookie stamping into existing
`proxy.ts` (Next.js 16 rejects both middleware.ts and proxy.ts)

---------

Co-authored-by: pamelachia <26612111+pamelachia@users.noreply.github.com>
Co-authored-by: Pamela Chia <pamelachiamayyee@gmail.com>
2026-02-23 13:31:39 -08:00