mirror of
https://github.com/supabase/supabase.git
synced 2026-10-06 18:05:11 +03:00
Part 2 of a 5-PR stack on `apps/docs/content/guides/database/tables.mdx`. Builds on #50021. ## Problem Four structural problems, all covered by the Guides section of CONTRIBUTING. **Views was a second topic living inside a guide about tables.** Roughly 180 lines, its own subsections four levels deep, sharing nothing with the tables above it beyond the word "table". **The page didn't say what it was for.** It opened with three paragraphs and a sample table before a reader could tell whether the page matched their goal. CONTRIBUTING asks a guide to begin with a sentence declaring its intent. **The top level mixed information types.** It was a flat list of every task, so "Schemas" and "Primary keys" sat beside "Creating tables" and background interrupted the action path. **Reference material interrupted the procedure.** A 44-row data type table sat between "Creating tables" and "Loading data", so a reader following the action path walked through it. Plus a duplicate video: the Dashboard tab under "Joining tables with foreign keys" embedded the same YouTube ID that frontmatter already serves as the table of contents video. ## Solution Moves and regrouping. - **Views moves to its own page**, `guides/database/views`, with its headings promoted one level and the two view-related links from Resources moved with it. - The page opens with an intent sentence, then a section outline, then a "What is a table?" section holding the definition and the spreadsheet comparison. - The remaining sections split into three groups by information type, ordered procedures, context, reference: **Creating and managing tables** holds creating, loading, and joining; **How tables are organized** holds primary keys, relationships, and schemas; **Reference** holds the data type table. - "Joining tables with foreign keys" held both classes, so it splits. The steps keep the heading and stay in the procedures group. The concept, what relational means and the diagram showing it, becomes **Relationships between tables** in the context group. The two cross-reference each other. - The duplicate video goes, and with the Dashboard tab empty the surrounding `Tabs` wrapper goes too. ## Anchors **Every heading keeps its text, so every anchor keeps its slug.** Demoting a heading changes its level, not its anchor. That matters because the inbound links are mostly outside `apps/docs`: Studio deep-links to `#data-types` from three components and `#primary-keys` from two, and `apps/www` links to `#creating-tables` and `#joining-tables-with-foreign-keys`. `#views` is the one exception, since that content left the page. Its single inbound link, in `guides/ai/engineering-for-scale.mdx`, now points at the new page, and both `NavigationMenu.constants.ts` entries are updated: the existing item becomes "Managing tables and data" and a "Views" item sits beside it. ## One deletion that isn't a move The "Columns" heading and its one sentence, "You must define the data type when you create a column." The heading held only the two subsections that moved out, and the sentence repeats a line 50 lines above it. ## Deferred Reordering "View security" behind an access-control foundation. That move only reads correctly once the foundation exists, so it travels with that content in #50024. ## Manual testing Preview: https://docs-git-docs-tables-structure-supabase.vercel.app/docs/guides/database/tables 1. Open the preview. The page opens with its intent, then a four-entry outline, then "What is a table?". Each outline link resolves, and the three groups below read as procedures, then context, then reference. 2. Open `#data-types`, `#primary-keys`, `#creating-tables`, and `#joining-tables-with-foreign-keys` on the preview. All four still land on their sections. 3. Open https://docs-git-docs-tables-structure-supabase.vercel.app/docs/guides/database/views. The new page renders, and "Views" appears in the sidebar beside "Managing tables and data". 4. Run `pnpm build:guides-markdown` from `apps/docs`. It generates 782 files, one more than before. Discard the change to `public/markdown/manifest.json`, which the repo commits as `[]`. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Documentation** - Added a dedicated guide covering Postgres views, including creation, querying, security options, and materialized views. - Reorganized the Tables and data guide with clearer sections, navigation links, table organization details, and reference information. - Updated the many-to-many example to display SQL directly. - Split database navigation into separate “Managing tables and data” and “Views” entries. - Added a PostgreSQL log configuration entry and a C# client reference link. - Updated documentation links to point to the new Views guide. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
150 lines
5.7 KiB
Plaintext
150 lines
5.7 KiB
Plaintext
---
|
|
id: 'ai-engineering-for-scale'
|
|
title: 'Engineering for Scale'
|
|
description: 'Building an enterprise-grade vector architecture'
|
|
subtitle: 'Building an enterprise-grade vector architecture.'
|
|
sidebar_label: 'Engineering for Scale'
|
|
---
|
|
|
|
Content sources for vectors can be extremely large. As you grow you should run your Vector workloads across several secondary databases (sometimes called "pods"), which allows each collection to scale independently.
|
|
|
|
## Small workloads [#simple-workloads]
|
|
|
|
For small workloads, you can typically store your data in a single database.
|
|
|
|
If you've used [Vecs](/docs/guides/ai/vecs-python-client) to create 3 different collections, you can expose collections to your web or mobile application using [views](/docs/guides/database/views):
|
|
|
|
The diagram below shows a single database holding the three vector collections of `docs`, `posts`, and `images`. Each are exposed to your application through a view.
|
|
|
|
<Image
|
|
alt="Architecture diagram: a single Supabase database holding three vector collections (docs, posts, and images), each exposed to the application through a view."
|
|
src={{
|
|
light: '/docs/img/ai/scaling/engineering-for-scale--single-database--light.png',
|
|
dark: '/docs/img/ai/scaling/engineering-for-scale--single-database--dark.png',
|
|
}}
|
|
width={1600}
|
|
height={1145}
|
|
/>
|
|
|
|
For example, with 3 collections, called `docs`, `posts`, and `images`, we could expose the "docs" inside the public schema like this:
|
|
|
|
```sql
|
|
create view public.docs as
|
|
select
|
|
id,
|
|
embedding,
|
|
metadata, # Expose the metadata as JSON
|
|
(metadata->>'url')::text as url # Extract the URL as a string
|
|
from vector
|
|
```
|
|
|
|
You can then use any of the client libraries to access your collections within your applications:
|
|
|
|
{/* prettier-ignore */}
|
|
```js
|
|
const { data, error } = await supabase
|
|
.from('docs')
|
|
.select('id, embedding, metadata')
|
|
.eq('url', '/hello-world')
|
|
```
|
|
|
|
## Enterprise workloads
|
|
|
|
As you move into production, we recommend splitting your collections into separate projects. This is because it allows your vector stores to scale independently of your production data. Vectors typically grow faster than operational data, and they have different resource requirements. Running them on separate databases removes the single-point-of-failure.
|
|
|
|
The diagram below shows a primary database alongside separate secondary "pod" databases, each holding its own vector collection so they can scale independently.
|
|
|
|
<Image
|
|
alt="Architecture diagram: a primary database alongside separate secondary 'pod' databases, each holding its own vector collection so collections can scale independently."
|
|
src={{
|
|
light: '/docs/img/ai/scaling/engineering-for-scale--with-secondaries--light.png',
|
|
dark: '/docs/img/ai/scaling/engineering-for-scale--with-secondaries--dark.png',
|
|
}}
|
|
width={1600}
|
|
height={1641}
|
|
/>
|
|
|
|
You can use as many secondary databases as you need to manage your collections. With this architecture, you have 2 options for accessing collections within your application:
|
|
|
|
1. Query the collections directly using Vecs.
|
|
2. Access the collections from your Primary database through a Wrapper.
|
|
|
|
You can use both of these in tandem to suit your use-case. We recommend option `1` wherever possible, as it offers the most scalability.
|
|
|
|
### Query collections using Vecs
|
|
|
|
Vecs provides methods for querying collections, either using a [cosine similarity function](https://supabase.github.io/vecs/api/#basic) or with [metadata filtering](https://supabase.github.io/vecs/api/#metadata-filtering).
|
|
|
|
```python
|
|
# cosine similarity
|
|
docs.query(query_vector=[0.4,0.5,0.6], limit=5)
|
|
|
|
# metadata filtering
|
|
docs.query(
|
|
query_vector=[0.4,0.5,0.6],
|
|
limit=5,
|
|
filters={"year": {"$eq": 2012}}, # metadata filters
|
|
)
|
|
```
|
|
|
|
### Accessing external collections using Wrappers
|
|
|
|
Supabase supports [Foreign Data Wrappers](/blog/postgres-foreign-data-wrappers-rust). Wrappers allow you to connect two databases together so that you can query them over the network.
|
|
|
|
This involves 2 steps: connecting to your remote database from the primary and creating a Foreign Table.
|
|
|
|
#### Connecting your remote database
|
|
|
|
Inside your Primary database we need to provide the credentials to access the secondary database:
|
|
|
|
```sql
|
|
create extension postgres_fdw;
|
|
|
|
create server docs_server
|
|
foreign data wrapper postgres_fdw
|
|
options (host 'db.xxx.supabase.co', port '5432', dbname 'postgres');
|
|
|
|
create user mapping for docs_user
|
|
server docs_server
|
|
options (user 'postgres', password 'password');
|
|
```
|
|
|
|
#### Create a foreign table
|
|
|
|
We can now create a foreign table to access the data in our secondary project.
|
|
|
|
```sql
|
|
create foreign table docs (
|
|
id text not null,
|
|
embedding extensions.vector(384),
|
|
metadata jsonb,
|
|
url text
|
|
)
|
|
server docs_server
|
|
options (schema_name 'public', table_name 'docs');
|
|
```
|
|
|
|
This looks very similar to our View example above, and you can continue to use the client libraries to access your collections through the foreign table:
|
|
|
|
{/* prettier-ignore */}
|
|
```js
|
|
const { data, error } = await supabase
|
|
.from('docs')
|
|
.select('id, embedding, metadata')
|
|
.eq('url', '/hello-world')
|
|
```
|
|
|
|
### Enterprise architecture
|
|
|
|
This diagram below provides an example architecture that allows you to access the collections either with our client libraries or using Vecs. You can add as many secondary databases as you need (in this example we only show one):
|
|
|
|
<Image
|
|
alt="Enterprise architecture diagram: an application accessing vector collections in multiple secondary databases, either directly via Vecs or through the primary database using Foreign Data Wrappers."
|
|
src={{
|
|
light: '/docs/img/ai/scaling/engineering-for-scale--multi-database--light.png',
|
|
dark: '/docs/img/ai/scaling/engineering-for-scale--multi-database--dark.png',
|
|
}}
|
|
width={1600}
|
|
height={1754}
|
|
/>
|