mirror of
https://github.com/supabase/supabase.git
synced 2026-10-09 11:25:06 +03:00
Closes DOCS-1057 Contributes to DOCS-1052 ## I have read the [CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md) file. YES ## Problem We have hundreds of MDX lint warnings in our docs going against style best practices. ## Solution Remove and replace in context the following: - PostgreSQL. There was only one. There was concern about exceptions, but I found none. - Just - Quickly - Actually ### What changed Edits follow the [Google developer documentation style guide](https://developers.google.com/style): concise, direct, active voice. The flagged words were removed when the sentence still read well, or replaced when meaning needed to be preserved. ### Common patterns | Flagged word | Approach | Example | |---|---|---| | **just** (filler) | Removed | "you just installed" → "you installed" | | **just** (limiting) | **only** | "just one row" → "only one row" | | **just like** | **like** / **the same as** | "function just like regular users" → "function like regular users" | | **not just** | **not only** | "not just errors" → "not only errors" | | **quickly** (performance) | **efficiently** or removed | "find rows quickly" → "find rows efficiently" | | **quickly** (time) | **soon** / **rapidly** / removed | "expires too quickly" → "expires too soon" | | **actually** (filler) | Removed | "actually execute" → "execute"; "is actually the most common" → "is the most common" | ## Tophatting 1. See the diff. 2. See that content continues to make sense in context. 3. Locally, `cd apps/docs` and run `pnpm run lint:mdx`. 4. Search for "just," "actually," "quickly", and "PostgreSQL" and see there are 0 warnings. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Updated wording across quickstarts, guides, and troubleshooting articles for grammar, clarity, and consistent step-by-step phrasing. * Clarified key concepts including Row Level Security policy evaluation across Supabase products, deferred foreign key constraint behavior, and when `EXPLAIN ANALYZE` executes queries (and related side effects). * Refined several troubleshooting instructions and added guidance to cap log payload size to reduce billed Logs Ingest volume. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com> Co-authored-by: Nik Richers <nrichers@gmail.com> Co-authored-by: Chris Chinchilla <chris.ward@supabase.io>
140 lines
6.3 KiB
Plaintext
140 lines
6.3 KiB
Plaintext
---
|
|
id: 'function-ai-models'
|
|
title: 'Semantic Search'
|
|
description: 'Semantic Search with pgvector and Supabase Edge Functions'
|
|
subtitle: 'Semantic Search with pgvector and Supabase Edge Functions'
|
|
tocVideo: 'w4Rr_1whU-U'
|
|
---
|
|
|
|
[Semantic search](/docs/guides/ai/semantic-search) interprets the meaning behind user queries rather than exact [keywords](/docs/guides/ai/keyword-search). It uses machine learning to capture the intent and context behind the query, handling language nuances like synonyms, phrasing variations, and word relationships.
|
|
|
|
Since Supabase Edge Runtime [v1.36.0](https://github.com/supabase/edge-runtime/releases/tag/v1.36.0) you can run the [`gte-small` model](https://huggingface.co/Supabase/gte-small) natively within Supabase Edge Functions without any external dependencies! This allows you to generate text embeddings without calling any external APIs!
|
|
|
|
In this tutorial you're implementing three parts:
|
|
|
|
1. A [`generate-embedding`](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/functions/generate-embedding/index.ts) database webhook edge function which generates embeddings when a content row is added (or updated) in the [`public.embeddings`](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/migrations/20240408072601_embeddings.sql) table.
|
|
2. A [`query_embeddings` Postgres function](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/migrations/20240410031515_vector-search.sql) which allows us to perform similarity search from an Edge Function via [Remote Procedure Call (RPC)](/docs/guides/database/functions?language=js).
|
|
3. A [`search` edge function](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/functions/search/index.ts) which generates the embedding for the search term, performs the similarity search via RPC function call, and returns the result.
|
|
|
|
You can find the complete example code on [GitHub](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions)
|
|
|
|
### Create the database table and webhook
|
|
|
|
Given the [following table definition](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/migrations/20240408072601_embeddings.sql):
|
|
|
|
```sql
|
|
create extension if not exists vector with schema extensions;
|
|
|
|
create table embeddings (
|
|
id bigint primary key generated always as identity,
|
|
content text not null,
|
|
embedding extensions.vector (384)
|
|
);
|
|
alter table embeddings enable row level security;
|
|
|
|
create index on embeddings using hnsw (embedding vector_ip_ops);
|
|
```
|
|
|
|
You can deploy the [following edge function](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/functions/generate-embedding/index.ts) as a [database webhook](/docs/guides/database/webhooks) to generate the embeddings for any text content inserted into the table:
|
|
|
|
```ts
|
|
import { withSupabase } from 'npm:@supabase/server@^1'
|
|
|
|
const model = new Supabase.ai.Session('gte-small')
|
|
|
|
// Triggered by a Database Webhook, which authenticates with a secret key.
|
|
// Deploy with `verify_jwt = false`.
|
|
export default {
|
|
fetch: withSupabase({ auth: 'secret' }, async (req, ctx) => {
|
|
const payload: WebhookPayload = await req.json()
|
|
const { content, id } = payload.record
|
|
|
|
// Generate embedding.
|
|
const embedding = await model.run(content, {
|
|
mean_pool: true,
|
|
normalize: true,
|
|
})
|
|
|
|
// Store in database.
|
|
const { error } = await ctx.supabaseAdmin
|
|
.from('embeddings')
|
|
.update({ embedding: JSON.stringify(embedding) })
|
|
.eq('id', id)
|
|
if (error) console.warn(error.message)
|
|
|
|
return Response.json({ ok: true })
|
|
}),
|
|
}
|
|
```
|
|
|
|
## Create a Database Function and RPC
|
|
|
|
With the embeddings now stored in your Postgres database table, you can query them from Supabase Edge Functions by using [Remote Procedure Calls (RPC)](/docs/guides/database/functions?language=js).
|
|
|
|
Given the [following Postgres Function](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/migrations/20240410031515_vector-search.sql):
|
|
|
|
```sql
|
|
-- Matches document sections using vector similarity search on embeddings
|
|
--
|
|
-- Returns a setof embeddings so that we can use PostgREST resource embeddings (joins with other tables)
|
|
-- Additional filtering like limits can be chained to this function call
|
|
create or replace function query_embeddings(embedding extensions.vector(384), match_threshold float)
|
|
returns setof embeddings
|
|
language plpgsql
|
|
as $$
|
|
begin
|
|
return query
|
|
select *
|
|
from embeddings
|
|
|
|
-- The inner product is negative, so we negate match_threshold
|
|
where embeddings.embedding <#> embedding < -match_threshold
|
|
|
|
-- Our embeddings are normalized to length 1, so cosine similarity
|
|
-- and inner product will produce the same query results.
|
|
-- Using inner product which can be computed faster.
|
|
--
|
|
-- For the different distance functions, see https://github.com/pgvector/pgvector
|
|
order by embeddings.embedding <#> embedding;
|
|
end;
|
|
$$;
|
|
```
|
|
|
|
## Query vectors in Supabase Edge Functions
|
|
|
|
You can use `supabase-js` to first generate the embedding for the search term and then invoke the Postgres function to find the relevant results from your stored embeddings, right from your [Supabase Edge Function](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/functions/search/index.ts):
|
|
|
|
```ts
|
|
import { withSupabase } from 'npm:@supabase/server@^1'
|
|
|
|
const model = new Supabase.ai.Session('gte-small')
|
|
|
|
export default {
|
|
fetch: withSupabase({ auth: 'user' }, async (req, ctx) => {
|
|
const { search } = await req.json()
|
|
if (!search) return Response.json({ error: 'Please provide a search param!' }, { status: 400 })
|
|
// Generate embedding for search term.
|
|
const embedding = await model.run(search, {
|
|
mean_pool: true,
|
|
normalize: true,
|
|
})
|
|
|
|
// Query embeddings.
|
|
const { data: result, error } = await ctx.supabase
|
|
.rpc('query_embeddings', {
|
|
embedding,
|
|
match_threshold: 0.8,
|
|
})
|
|
.select('content')
|
|
.limit(3)
|
|
if (error) {
|
|
return Response.json({ error: error.message }, { status: 500 })
|
|
}
|
|
|
|
return Response.json({ search, result })
|
|
}),
|
|
}
|
|
```
|
|
|
|
You now have AI powered semantic search set up without any external dependencies! All you need: you, pgvector, and Supabase Edge Functions!
|