mirror of
https://github.com/supabase/supabase.git
synced 2026-10-06 09:55:06 +03:00
Closes [DOCS-1289](https://linear.app/supabase/issue/DOCS-1289/get-the-linter-to-fix-what-it-flags-or-retirereplace-the-linter) Stacked on #50600, which points contributors at the authoring skills. Merge that one first. ## Problem Contributors experienced friction with the linter. They felt nickle and dimed for tiny nits and felt detracted from the work itself. PRs would become noisy with tiny one-word suggestions. Additionally, our homegrown linter is not very intelligent, causing frequent overrides. ## Solution This removes the linter entirely in favor of directing contributors to use SKILLS instead. The removal entails... - **CI.** Delete the three `docs_lint` workflows: the PR check, the external-PR comment companion, and the nightly `--fix` bot. Drop the stale `zizmor.yml` ignore entry for the deleted workflow. - **Tooling.** Delete `supa-mdx-lint.config.toml` and the 14 rule files. Drop the `lint:mdx` script and the `@supabase/supa-mdx-lint` dependency from docs, learn, and ui-library, and regenerate the lockfile. - **Content.** Remove the 181 directives. A separate commit carries Prettier's reformatting of the tables and blank lines those comments had suppressed, so the deletion commit stays readable. No prose changes. - **Style guide.** The word list states each rule directly instead of describing what the linter flagged. Every term survives, including the phrase groups that mirrored `Rule004ExcludeWords`. - **Skills.** `write-the-docs`, `edit-the-docs`, and `review-the-docs` drop `pnpm lint:mdx` from their self-review commands and check the word list directly. `ask-the-docs`'s CI reference drops both workflows. ## Manual testing 1. Run `git grep -i supa-mdx-lint -- . ':!pnpm-lock.yaml'`. No matches. 2. Run `pnpm install --frozen-lockfile --lockfile-only`. It passes, so the lockfile matches the three trimmed manifests. 3. Run `git diff master...HEAD --name-only --diff-filter=ACMR | grep -E '\.(md|mdx)$' | xargs npx prettier --config prettier.config.mjs --check`. All changed markdown passes. 4. Open the [reformatted filter table](https://docs-git-docs-retire-mdx-linter-supabase.vercel.app/docs/guides/observability/logs#filter-events) on the preview and compare it with [production](https://supabase.com/docs/guides/observability/logs#filter-events). The table renders the same. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Documentation** * Documentation guidance now uses manual prose and terminology review with the shared word list. * Clarified storage configuration and common Realtime channel mistakes. * Improved table formatting, text wrapping, and selected reference links. * Updated documentation authoring and review guidance. * **Chores** * Retired automated MDX linting from workflows and local validation commands. * Removed lint-suppression markers throughout documentation without changing instructions. * Added targeted documentation review guidance for pull requests. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
69 lines
3.9 KiB
Plaintext
69 lines
3.9 KiB
Plaintext
---
|
|
id: 'ai-concepts'
|
|
title: 'Concepts'
|
|
description: 'Learn about embeddings within AI and vector applications.'
|
|
sidebar_label: 'Concepts'
|
|
---
|
|
|
|
Embeddings are core to many AI and vector applications. This guide covers these concepts. If you prefer to get started right away, see our guide on [Generating Embeddings](/docs/guides/ai/quickstarts/generate-text-embeddings).
|
|
|
|
## What are embeddings?
|
|
|
|
Embeddings capture the "relatedness" of text, images, video, or other types of information. This relatedness is most commonly used for:
|
|
|
|
- **Search:** how similar is a search term to a body of text?
|
|
- **Recommendations:** how similar are two products?
|
|
- **Classifications:** how do we categorize a body of text?
|
|
- **Clustering:** how do we identify trends?
|
|
|
|
The following example uses text embeddings. Given three phrases:
|
|
|
|
1. "The cat chases the mouse"
|
|
2. "The kitten hunts rodents"
|
|
3. "I like ham sandwiches"
|
|
|
|
Your job is to group phrases with similar meaning. If you are a human, this should be obvious. Phrases 1 and 2 are almost identical, while phrase 3 has a completely different meaning.
|
|
|
|
Although phrases 1 and 2 are similar, they share no common vocabulary (besides "the"). Yet their meanings are nearly identical. How can we teach a computer that these are the same?
|
|
|
|
## Human language
|
|
|
|
Humans use words and symbols to communicate language. But words in isolation are mostly meaningless - we need to draw from shared knowledge & experience in order to make sense of them. The phrase “You should Google it” only makes sense if you know that Google is a search engine and that people have been using it as a verb.
|
|
|
|
In the same way, we need to train a neural network model to understand human language. An effective model should be trained on millions of different examples to understand what each word, phrase, sentence, or paragraph could mean in different contexts.
|
|
|
|
So how does this relate to embeddings?
|
|
|
|
## How do embeddings work?
|
|
|
|
Embeddings compress discrete information (words & symbols) into distributed continuous-valued data (vectors). If we took our phrases from before and plot them on a chart, it might look something like this:
|
|
|
|
The chart below plots example phrases as points. Phrases with similar meanings sit close together, and unrelated phrases sit far apart.
|
|
|
|
<Image
|
|
src="/docs/img/ai/vector-similarity.png"
|
|
alt="A two-dimensional chart plotting example phrases as points, where phrases with similar meanings sit close together and unrelated phrases sit far apart."
|
|
width="640"
|
|
height="640"
|
|
/>
|
|
|
|
Phrases 1 and 2 would be plotted close to each other, since their meanings are similar. We would expect phrase 3 to live somewhere far away since it isn't related. If we had a fourth phrase, “Sally ate Swiss cheese”, this might exist somewhere between phrase 3 (cheese can go on sandwiches) and phrase 1 (mice like Swiss cheese).
|
|
|
|
In this example we only have 2 dimensions: the X and Y axis. In reality, we would need many more dimensions to effectively capture the complexities of human language.
|
|
|
|
## Using embeddings
|
|
|
|
Compared to our 2-dimensional example above, most embedding models will output many more dimensions. For example the open source [`gte-small`](https://huggingface.co/Supabase/gte-small) model outputs 384 dimensions.
|
|
|
|
Why is this useful? Once we have generated embeddings on multiple texts, it is trivial to calculate how similar they are using vector math operations like cosine distance. A common use case for this is search. Your process might look something like this:
|
|
|
|
1. Pre-process your knowledge base and generate embeddings for each page
|
|
2. Store your embeddings to be referenced later
|
|
3. Build a search page that prompts your user for input
|
|
4. Take user's input, generate a one-time embedding, then perform a similarity search against your pre-processed embeddings.
|
|
5. Return the most similar pages to the user
|
|
|
|
## See also
|
|
|
|
- [Structured and Unstructured embeddings](/docs/guides/ai/structured-unstructured)
|