Files
supabase/apps/www/content/md/modules/vector.md
T
Pamela Chia 9113c2ba04 feat: markdown alternate tags + llms.txt cleanup (#48287)
The June/July marketing redesign (#47271, #47228) rebuilt the homepage
and product pages off the Pages Router, silently dropping their `<link
rel="alternate" type="text/markdown">` head tags, and llms-full.txt has
been accidentally embedding every blog/customer/event page via an
`MD_CONTENT` spread. I restored the tags behind a shared helper, added a
CI drift test so a future redesign can't drop them silently again, and
trimmed both llms files to the agreed docs-index shape.

**Changed:**

- **Markdown siblings advertised again**: homepage, the 5 product pages,
pricing, and blog emit absolute `.md` alternate URLs via a new
`mdAlternates(slug)` helper (the one documented consumer of the tag
parses it from `<head>` and fetches the `.md` sibling, so tags must
point at the sibling, never the page itself).
- **Drift test**: a vitest file walks `content/md/**` and asserts every
markdown-served slug's page wires the helper (or is covered by the Pages
Router `_app.tsx` mechanism, whose alternate-link wiring the test also
asserts directly so removing it fails CI too). Source-level assertions
by design: page modules can't be imported under www's vitest config.
Fails correctly when wiring is removed (verified by hiding a page and by
altering the `_app.tsx` tag).
- **Vector orphan fixed**: `content/md/vector.md` moved to
`modules/vector` matching the live route (the page previously had no
negotiation or tag, and `/modules/vector.md` 404'd); `/vector.md` now
308s to `/modules/vector.md` and the legacy `/llms/vector.txt` redirect
no longer chains.
- **llms.txt + llms-full.txt**: the `## Product Overview` sections are
gone from both, each keeps a `## Pricing` section. This deletes the
hand-maintained links array (a drift trap) and fixes the accidental
~470-page embed, shrinking llms-full.txt from ~9.8MB to ~4.9MB and
dropping the 4.1MB generated content module from that route's serverless
bundle.

**Note:** this PR is scoped to apps/www only. The docs side
(troubleshooting pages and the rest of the docs surface) is handled
separately through a consolidated manifest-gated mechanism; an earlier
troubleshooting-tag commit was reverted out of this branch to keep the
scopes clean.

<details>
<summary>Why alternate tags matter (background)</summary>

Agents ingest markdown far more efficiently than our rendered HTML: a
fraction of the tokens and no extraction step. Since #47770 removed
UA-based serving (UA sniffing broke a major AI app's fetcher and
poisoned CDN caches), markdown is served only on explicit request: a
`.md` suffix URL, an `Accept: text/markdown` header, or llms.txt. That's
the right serving model, but it makes the markdown twin invisible to any
agent that doesn't already know our URL convention, and the major AI
fetchers send browser/wildcard Accept headers, so bare URLs hand them
HTML.

The `<link rel="alternate" type="text/markdown">` head tag is the
standards-based advertisement of the sibling. It has a documented
consumer today: an agent CLI that parses the tag from `<head>` and then
fetches the `.md` sibling, which is also why the tag must point at a
real sibling URL and never at the page itself. Peer docs sites ship this
tag as table stakes. These www pages used to carry it until the
June/July marketing redesign silently dropped it; the drift test in this
PR turns that regression class into a CI failure.

</details>

## To test

Tested locally (www + docs dev servers):
- [x] `/llms.txt` renders `## Documentation` + single-link `## Pricing`,
no Product Overview
- [x] `/llms-full.txt` renders `# Supabase` → `## Pricing` → `##
Documentation`, no Product Overview, ~4.9MB
- [x] Full www suite: 6 files / 71 tests green; drift test fails
correctly when a page is removed or the `_app.tsx` wiring is altered
- [x] `generateMdContent.mjs` emits `modules/vector`, bare `vector` slug
gone

On the Vercel preview (browser-verified with Playwright):
- [x] Alternate tag present on `/`, `/auth`, `/database`, `/storage`,
`/edge-functions`, `/realtime`, `/pricing`, and a blog post: exactly one
tag each, href = preview origin + `.md` sibling
- [x] `/vector.md` → 308 → `/modules/vector.md`, renders as markdown (`#
Supabase Vector`)
- [x] `/llms.txt` shows single-link `## Pricing`, no Product Overview
- [x] Coverage sweep: all 482 `MD_PAGES` slugs + changelog index/entry
curled on the preview; 471 pages carry exactly one tag, all `.md`
siblings 200 as `text/markdown`. The 11 misses are legacy blog slugs
whose HTML 308-redirects away (stale `MD_PAGES` entries predating this
PR, no head to tag; follow-up tracked in Linear)

Post-merge prod:
- [ ] Full llms.txt link sweep (every linked URL 200s; previews can't
cover the docs-hosted links)

## Linear

- fixes GROWTH-1013


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added Markdown alternate links across key product, pricing, blog, and
troubleshooting pages.
* Added Supabase Vector documentation covering features, use cases,
workflows, and technical details.
* Updated AI-focused documentation indexes with dedicated pricing
content.
  * Added redirects for updated Vector documentation URLs.

* **Tests**
* Added coverage to verify Markdown documentation links stay aligned
with available pages.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-07-27 17:06:16 +08:00

2.3 KiB

Supabase Vector

Store, index, and query vector embeddings in Postgres with pgvector.

Supabase Vector is an AI toolkit that lets you store vector embeddings alongside your transactional data in the same Postgres database. Powered by the pgvector extension, it eliminates the need for a separate vector database while providing production-grade similarity search.

Key Features

  • pgvector integration: store, index, and query vector embeddings directly in Postgres
  • Co-located data: vector embeddings live alongside your relational data, joined with standard SQL
  • Multiple index types: IVFFlat and HNSW indexes for fast approximate nearest neighbor search
  • Distance metrics: cosine distance, L2 (Euclidean) distance, max inner product
  • Metadata filtering: filter similarity queries by any column or JSONB metadata
  • Python client (vecs): dedicated Python library for managing collections, upserting vectors, and querying
  • LLM integrations: works with OpenAI, Hugging Face, Amazon SageMaker, LangChain, and more
  • Edge Functions: generate embeddings using open source models directly in Edge Functions
  • Self-hostable: run the full stack on your own infrastructure
  • SOC2 Type 2 compliant: enterprise-grade security

Common Use Cases

  • Semantic search over documents, knowledge bases, or support tickets
  • Retrieval-Augmented Generation (RAG) for LLM applications
  • Image similarity detection
  • Recommendation engines
  • Auto-tagging and content classification
  • ChatGPT plugins with long-term memory

How It Works

  1. Generate embeddings using any model (OpenAI, Hugging Face, Cohere, etc.)
  2. Store embeddings in a Postgres table with a vector column
  3. Create an HNSW or IVFFlat index for fast similarity search
  4. Query with <=> (cosine), <-> (L2), or <#> (inner product) operators
  5. Combine with standard SQL: JOIN, WHERE, GROUP BY for hybrid queries

Technical Details

  • Extension: pgvector (open source)
  • Max dimensions: 2,000 (HNSW), 16,000+ (flat)
  • Index types: HNSW (recommended), IVFFlat
  • Scaling: same compute scaling as your Supabase database (Micro to 16XL)
  • Backups: automatic daily + PITR available