Files
supabase/apps/docs/content/guides/functions/examples/semantic-search.mdx
3dffdefd6e fix(docs) Resolve 196 mdx lint warnings for just, quickly, actually, PostgreSQL (#47358)
Closes DOCS-1057
Contributes to DOCS-1052

## I have read the
[CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md)
file.

YES

## Problem

We have hundreds of MDX lint warnings in our docs going against style
best practices.

## Solution

Remove and replace in context the following:

- PostgreSQL. There was only one. There was concern about exceptions,
but I found none.
- Just
- Quickly
- Actually

### What changed

Edits follow the [Google developer documentation style
guide](https://developers.google.com/style): concise, direct, active
voice. The flagged words were removed when the sentence still read well,
or replaced when meaning needed to be preserved.

### Common patterns

| Flagged word | Approach | Example |
|---|---|---|
| **just** (filler) | Removed | "you just installed" → "you installed" |
| **just** (limiting) | **only** | "just one row" → "only one row" |
| **just like** | **like** / **the same as** | "function just like
regular users" → "function like regular users" |
| **not just** | **not only** | "not just errors" → "not only errors" |
| **quickly** (performance) | **efficiently** or removed | "find rows
quickly" → "find rows efficiently" |
| **quickly** (time) | **soon** / **rapidly** / removed | "expires too
quickly" → "expires too soon" |
| **actually** (filler) | Removed | "actually execute" → "execute"; "is
actually the most common" → "is the most common" |

## Tophatting

1. See the diff.
2. See that content continues to make sense in context.
3. Locally, `cd apps/docs` and run `pnpm run lint:mdx`.
4. Search for "just," "actually," "quickly", and "PostgreSQL" and see
there are 0 warnings.




<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **Documentation**
* Updated wording across quickstarts, guides, and troubleshooting
articles for grammar, clarity, and consistent step-by-step phrasing.
* Clarified key concepts including Row Level Security policy evaluation
across Supabase products, deferred foreign key constraint behavior, and
when `EXPLAIN ANALYZE` executes queries (and related side effects).
* Refined several troubleshooting instructions and added guidance to cap
log payload size to reduce billed Logs Ingest volume.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: github-actions[bot] <41898282+github-actions[bot]@users.noreply.github.com>
Co-authored-by: Nik Richers <nrichers@gmail.com>
Co-authored-by: Chris Chinchilla <chris.ward@supabase.io>
2026-06-29 09:40:25 -07:00

140 lines
6.3 KiB
Plaintext

---
id: 'function-ai-models'
title: 'Semantic Search'
description: 'Semantic Search with pgvector and Supabase Edge Functions'
subtitle: 'Semantic Search with pgvector and Supabase Edge Functions'
tocVideo: 'w4Rr_1whU-U'
---
[Semantic search](/docs/guides/ai/semantic-search) interprets the meaning behind user queries rather than exact [keywords](/docs/guides/ai/keyword-search). It uses machine learning to capture the intent and context behind the query, handling language nuances like synonyms, phrasing variations, and word relationships.
Since Supabase Edge Runtime [v1.36.0](https://github.com/supabase/edge-runtime/releases/tag/v1.36.0) you can run the [`gte-small` model](https://huggingface.co/Supabase/gte-small) natively within Supabase Edge Functions without any external dependencies! This allows you to generate text embeddings without calling any external APIs!
In this tutorial you're implementing three parts:
1. A [`generate-embedding`](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/functions/generate-embedding/index.ts) database webhook edge function which generates embeddings when a content row is added (or updated) in the [`public.embeddings`](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/migrations/20240408072601_embeddings.sql) table.
2. A [`query_embeddings` Postgres function](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/migrations/20240410031515_vector-search.sql) which allows us to perform similarity search from an Edge Function via [Remote Procedure Call (RPC)](/docs/guides/database/functions?language=js).
3. A [`search` edge function](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions/supabase/functions/search/index.ts) which generates the embedding for the search term, performs the similarity search via RPC function call, and returns the result.
You can find the complete example code on [GitHub](https://github.com/supabase/supabase/tree/master/examples/ai/edge-functions)
### Create the database table and webhook
Given the [following table definition](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/migrations/20240408072601_embeddings.sql):
```sql
create extension if not exists vector with schema extensions;
create table embeddings (
id bigint primary key generated always as identity,
content text not null,
embedding extensions.vector (384)
);
alter table embeddings enable row level security;
create index on embeddings using hnsw (embedding vector_ip_ops);
```
You can deploy the [following edge function](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/functions/generate-embedding/index.ts) as a [database webhook](/docs/guides/database/webhooks) to generate the embeddings for any text content inserted into the table:
```ts
import { withSupabase } from 'npm:@supabase/server@^1'
const model = new Supabase.ai.Session('gte-small')
// Triggered by a Database Webhook, which authenticates with a secret key.
// Deploy with `verify_jwt = false`.
export default {
fetch: withSupabase({ auth: 'secret' }, async (req, ctx) => {
const payload: WebhookPayload = await req.json()
const { content, id } = payload.record
// Generate embedding.
const embedding = await model.run(content, {
mean_pool: true,
normalize: true,
})
// Store in database.
const { error } = await ctx.supabaseAdmin
.from('embeddings')
.update({ embedding: JSON.stringify(embedding) })
.eq('id', id)
if (error) console.warn(error.message)
return Response.json({ ok: true })
}),
}
```
## Create a Database Function and RPC
With the embeddings now stored in your Postgres database table, you can query them from Supabase Edge Functions by using [Remote Procedure Calls (RPC)](/docs/guides/database/functions?language=js).
Given the [following Postgres Function](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/migrations/20240410031515_vector-search.sql):
```sql
-- Matches document sections using vector similarity search on embeddings
--
-- Returns a setof embeddings so that we can use PostgREST resource embeddings (joins with other tables)
-- Additional filtering like limits can be chained to this function call
create or replace function query_embeddings(embedding extensions.vector(384), match_threshold float)
returns setof embeddings
language plpgsql
as $$
begin
return query
select *
from embeddings
-- The inner product is negative, so we negate match_threshold
where embeddings.embedding <#> embedding < -match_threshold
-- Our embeddings are normalized to length 1, so cosine similarity
-- and inner product will produce the same query results.
-- Using inner product which can be computed faster.
--
-- For the different distance functions, see https://github.com/pgvector/pgvector
order by embeddings.embedding <#> embedding;
end;
$$;
```
## Query vectors in Supabase Edge Functions
You can use `supabase-js` to first generate the embedding for the search term and then invoke the Postgres function to find the relevant results from your stored embeddings, right from your [Supabase Edge Function](https://github.com/supabase/supabase/blob/master/examples/ai/edge-functions/supabase/functions/search/index.ts):
```ts
import { withSupabase } from 'npm:@supabase/server@^1'
const model = new Supabase.ai.Session('gte-small')
export default {
fetch: withSupabase({ auth: 'user' }, async (req, ctx) => {
const { search } = await req.json()
if (!search) return Response.json({ error: 'Please provide a search param!' }, { status: 400 })
// Generate embedding for search term.
const embedding = await model.run(search, {
mean_pool: true,
normalize: true,
})
// Query embeddings.
const { data: result, error } = await ctx.supabase
.rpc('query_embeddings', {
embedding,
match_threshold: 0.8,
})
.select('content')
.limit(3)
if (error) {
return Response.json({ error: error.message }, { status: 500 })
}
return Response.json({ search, result })
}),
}
```
You now have AI powered semantic search set up without any external dependencies! All you need: you, pgvector, and Supabase Edge Functions!