mirror of
https://github.com/supabase/supabase.git
synced 2026-10-06 09:55:06 +03:00
The June/July marketing redesign (#47271, #47228) rebuilt the homepage and product pages off the Pages Router, silently dropping their `<link rel="alternate" type="text/markdown">` head tags, and llms-full.txt has been accidentally embedding every blog/customer/event page via an `MD_CONTENT` spread. I restored the tags behind a shared helper, added a CI drift test so a future redesign can't drop them silently again, and trimmed both llms files to the agreed docs-index shape. **Changed:** - **Markdown siblings advertised again**: homepage, the 5 product pages, pricing, and blog emit absolute `.md` alternate URLs via a new `mdAlternates(slug)` helper (the one documented consumer of the tag parses it from `<head>` and fetches the `.md` sibling, so tags must point at the sibling, never the page itself). - **Drift test**: a vitest file walks `content/md/**` and asserts every markdown-served slug's page wires the helper (or is covered by the Pages Router `_app.tsx` mechanism, whose alternate-link wiring the test also asserts directly so removing it fails CI too). Source-level assertions by design: page modules can't be imported under www's vitest config. Fails correctly when wiring is removed (verified by hiding a page and by altering the `_app.tsx` tag). - **Vector orphan fixed**: `content/md/vector.md` moved to `modules/vector` matching the live route (the page previously had no negotiation or tag, and `/modules/vector.md` 404'd); `/vector.md` now 308s to `/modules/vector.md` and the legacy `/llms/vector.txt` redirect no longer chains. - **llms.txt + llms-full.txt**: the `## Product Overview` sections are gone from both, each keeps a `## Pricing` section. This deletes the hand-maintained links array (a drift trap) and fixes the accidental ~470-page embed, shrinking llms-full.txt from ~9.8MB to ~4.9MB and dropping the 4.1MB generated content module from that route's serverless bundle. **Note:** this PR is scoped to apps/www only. The docs side (troubleshooting pages and the rest of the docs surface) is handled separately through a consolidated manifest-gated mechanism; an earlier troubleshooting-tag commit was reverted out of this branch to keep the scopes clean. <details> <summary>Why alternate tags matter (background)</summary> Agents ingest markdown far more efficiently than our rendered HTML: a fraction of the tokens and no extraction step. Since #47770 removed UA-based serving (UA sniffing broke a major AI app's fetcher and poisoned CDN caches), markdown is served only on explicit request: a `.md` suffix URL, an `Accept: text/markdown` header, or llms.txt. That's the right serving model, but it makes the markdown twin invisible to any agent that doesn't already know our URL convention, and the major AI fetchers send browser/wildcard Accept headers, so bare URLs hand them HTML. The `<link rel="alternate" type="text/markdown">` head tag is the standards-based advertisement of the sibling. It has a documented consumer today: an agent CLI that parses the tag from `<head>` and then fetches the `.md` sibling, which is also why the tag must point at a real sibling URL and never at the page itself. Peer docs sites ship this tag as table stakes. These www pages used to carry it until the June/July marketing redesign silently dropped it; the drift test in this PR turns that regression class into a CI failure. </details> ## To test Tested locally (www + docs dev servers): - [x] `/llms.txt` renders `## Documentation` + single-link `## Pricing`, no Product Overview - [x] `/llms-full.txt` renders `# Supabase` → `## Pricing` → `## Documentation`, no Product Overview, ~4.9MB - [x] Full www suite: 6 files / 71 tests green; drift test fails correctly when a page is removed or the `_app.tsx` wiring is altered - [x] `generateMdContent.mjs` emits `modules/vector`, bare `vector` slug gone On the Vercel preview (browser-verified with Playwright): - [x] Alternate tag present on `/`, `/auth`, `/database`, `/storage`, `/edge-functions`, `/realtime`, `/pricing`, and a blog post: exactly one tag each, href = preview origin + `.md` sibling - [x] `/vector.md` → 308 → `/modules/vector.md`, renders as markdown (`# Supabase Vector`) - [x] `/llms.txt` shows single-link `## Pricing`, no Product Overview - [x] Coverage sweep: all 482 `MD_PAGES` slugs + changelog index/entry curled on the preview; 471 pages carry exactly one tag, all `.md` siblings 200 as `text/markdown`. The 11 misses are legacy blog slugs whose HTML 308-redirects away (stale `MD_PAGES` entries predating this PR, no head to tag; follow-up tracked in Linear) Post-merge prod: - [ ] Full llms.txt link sweep (every linked URL 200s; previews can't cover the docs-hosted links) ## Linear - fixes GROWTH-1013 <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added Markdown alternate links across key product, pricing, blog, and troubleshooting pages. * Added Supabase Vector documentation covering features, use cases, workflows, and technical details. * Updated AI-focused documentation indexes with dedicated pricing content. * Added redirects for updated Vector documentation URLs. * **Tests** * Added coverage to verify Markdown documentation links stay aligned with available pages. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
2.3 KiB
2.3 KiB
Supabase Vector
Store, index, and query vector embeddings in Postgres with pgvector.
Supabase Vector is an AI toolkit that lets you store vector embeddings alongside your transactional data in the same Postgres database. Powered by the pgvector extension, it eliminates the need for a separate vector database while providing production-grade similarity search.
Key Features
- pgvector integration: store, index, and query vector embeddings directly in Postgres
- Co-located data: vector embeddings live alongside your relational data, joined with standard SQL
- Multiple index types: IVFFlat and HNSW indexes for fast approximate nearest neighbor search
- Distance metrics: cosine distance, L2 (Euclidean) distance, max inner product
- Metadata filtering: filter similarity queries by any column or JSONB metadata
- Python client (vecs): dedicated Python library for managing collections, upserting vectors, and querying
- LLM integrations: works with OpenAI, Hugging Face, Amazon SageMaker, LangChain, and more
- Edge Functions: generate embeddings using open source models directly in Edge Functions
- Self-hostable: run the full stack on your own infrastructure
- SOC2 Type 2 compliant: enterprise-grade security
Common Use Cases
- Semantic search over documents, knowledge bases, or support tickets
- Retrieval-Augmented Generation (RAG) for LLM applications
- Image similarity detection
- Recommendation engines
- Auto-tagging and content classification
- ChatGPT plugins with long-term memory
How It Works
- Generate embeddings using any model (OpenAI, Hugging Face, Cohere, etc.)
- Store embeddings in a Postgres table with a
vectorcolumn - Create an HNSW or IVFFlat index for fast similarity search
- Query with
<=>(cosine),<->(L2), or<#>(inner product) operators - Combine with standard SQL: JOIN, WHERE, GROUP BY for hybrid queries
Technical Details
- Extension: pgvector (open source)
- Max dimensions: 2,000 (HNSW), 16,000+ (flat)
- Index types: HNSW (recommended), IVFFlat
- Scaling: same compute scaling as your Supabase database (Micro to 16XL)
- Backups: automatic daily + PITR available
Links
- Documentation: https://supabase.com/docs/guides/ai
- Python client: https://supabase.com/docs/guides/ai/vecs-python-client
- Dashboard: https://supabase.com/dashboard