mirror of
https://github.com/supabase/supabase.git
synced 2026-10-06 18:05:11 +03:00
## Problem For the new search, we will scan all files present within the public/markdown directory, extract text nodes and upsert a new Supabase table in a new project to do FTS type search. ## Solution In this PR: - A new script directory is created with files to solve all the steps described above. - Unit tests added for the fundamental bits of the script. - A new workflow file is added so the action runs after push every time content is altered, added or removed. <!-- ## Preview links If relevant, include links to changed pages for easy review access. Copy the preview base URL from the Vercel bot comment on this PR. Use the following table as an example template. | Site | Live | Preview | Search for | | -------------- | ------------------------------------------------------------------------- | ------------------------------------------------------------------------------------------------------------ | ----------------------------- | | WWW | [/blog/your-post](https://supabase.com/blog/your-post) | [/blog/your-post](https://zone-www-dot-com-git-branch-name-supabase.vercel.app/blog/your-post) | unique phrase from the change | | Docs | [/docs/guides/your-page](https://supabase.com/docs/guides/your-page) | [/docs/guides/your-page](https://docs-git-branch-name-supabase.vercel.app/docs/guides/your-page) | unique phrase from the change | | Studio | [/dashboard](https://supabase.com/dashboard) | [/dashboard](https://studio-git-branch-name-supabase.vercel.app/dashboard) | unique phrase from the change | | Design system | [/design-system](https://supabase.com/design-system) | [/design-system](https://design-system-git-branch-name-supabase.vercel.app/design-system) | unique phrase from the change | | UI library | [/library](https://supabase.com/library) | [/library](https://ui-library-git-branch-name-supabase.vercel.app/library) | unique phrase from the change | | Knowledge base | [/kb/guides/your-page](https://supabase.com/kb/guides/your-page) | [/kb/guides/your-page](https://kb-git-branch-name-supabase.vercel.app/kb/guides/your-page) | unique phrase from the change | --> <!-- ## Additional context Optionally add any other context or screenshots. --> ## Review instructions Sadly, testing this work is quite complex, but in case someone wants to: 1. Create a new Supabase project ton your personal space 1. Copy the id of the project and a secret key and add it to the new Search V2 environment variables as shown in the example file 1. Copy the content of the newly added `setup.sql` and run it on the SQL editor of your project. 1. Fetch this branch and on the search directory, run `search-v2:ingest` 1. Your project's table should have rows corresponding to the content from docs ## Checklist Check all before review: - [x] I have read [CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md) - [x] If I wrote a new docs topic or edited an existing topic, I used the `/write-the-docs` or `/edit-the-docs` skill, which references [WORD_LIST](https://github.com/supabase/supabase/blob/master/apps/docs/WORD_LIST.md) and the docs [CONTRIBUTING](https://github.com/supabase/supabase/blob/master/apps/docs/CONTRIBUTING.md) guide <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **New Features** - Documentation pages are now indexed for full-text search at the page and section level. - Search results can show the most relevant section from each page, with its title, heading, excerpt, and relevance score. - Search content is automatically refreshed when published Markdown documentation changes. - **Tests** - Added coverage for Markdown parsing, page structure, routing, and search-content generation. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
149 lines
4.5 KiB
TypeScript
149 lines
4.5 KiB
TypeScript
import { describe, expect, it } from 'vitest'
|
|
|
|
import {
|
|
extractExcerpt,
|
|
extractSections,
|
|
extractTitle,
|
|
nodeToText,
|
|
parseMarkdownAst,
|
|
parsePage,
|
|
} from './markdown.js'
|
|
|
|
const USERS_MD = `# Users
|
|
|
|
A **user** in Supabase Auth is someone with a user ID, stored in the Auth schema. You can restrict access via [RLS policies](https://example.com).
|
|
|
|
## Permanent and anonymous users
|
|
|
|
Supabase distinguishes between permanent and anonymous users.
|
|
|
|
- **Permanent users** are tied to PII.
|
|
- **Anonymous users** aren't tied to any identities.
|
|
|
|
## Inviting users
|
|
|
|
You can invite someone by email.
|
|
|
|
### Using the Dashboard
|
|
|
|
1. Go to **Authentication > Users**.
|
|
2. Click **Add user**.
|
|
|
|
### Using the Auth Admin API
|
|
|
|
Call \`inviteUserByEmail()\` from a server.
|
|
|
|
\`\`\`js
|
|
const x = 1
|
|
\`\`\`
|
|
`
|
|
|
|
describe('extractTitle', () => {
|
|
it('returns the first H1', () => {
|
|
expect(extractTitle(parseMarkdownAst(USERS_MD))).toBe('Users')
|
|
})
|
|
|
|
it('returns an empty string when there is no H1', () => {
|
|
expect(extractTitle(parseMarkdownAst('## Only an h2\n\nText.'))).toBe('')
|
|
})
|
|
|
|
it('strips inline formatting from the heading', () => {
|
|
expect(extractTitle(parseMarkdownAst('# The `code` **title**'))).toBe('The code title')
|
|
})
|
|
})
|
|
|
|
describe('extractExcerpt', () => {
|
|
it('returns the first paragraph as plain text', () => {
|
|
expect(extractExcerpt(parseMarkdownAst(USERS_MD))).toBe(
|
|
'A user in Supabase Auth is someone with a user ID, stored in the Auth schema. You can restrict access via RLS policies.'
|
|
)
|
|
})
|
|
|
|
it('skips YAML frontmatter', () => {
|
|
const md = `---\ntitle: Hello\n---\n\n# Title\n\nFirst paragraph.`
|
|
expect(extractExcerpt(parseMarkdownAst(md))).toBe('First paragraph.')
|
|
})
|
|
|
|
it('returns an empty string when there is no paragraph', () => {
|
|
expect(extractExcerpt(parseMarkdownAst('# Just a heading'))).toBe('')
|
|
})
|
|
})
|
|
|
|
describe('extractSections', () => {
|
|
const sections = extractSections(parseMarkdownAst(USERS_MD))
|
|
|
|
it('creates one section per heading', () => {
|
|
expect(sections.map((s) => s.heading)).toEqual([
|
|
'Users',
|
|
'Permanent and anonymous users',
|
|
'Inviting users',
|
|
'Using the Dashboard',
|
|
'Using the Auth Admin API',
|
|
])
|
|
})
|
|
|
|
it('records heading levels', () => {
|
|
expect(sections.map((s) => s.level)).toEqual([1, 2, 2, 3, 3])
|
|
})
|
|
|
|
it('builds the heading path from the H1 down to the section', () => {
|
|
expect(sections[1].headingPath).toEqual(['Users', 'Permanent and anonymous users'])
|
|
expect(sections[3].headingPath).toEqual(['Users', 'Inviting users', 'Using the Dashboard'])
|
|
})
|
|
|
|
it('resets the path when a sibling heading appears', () => {
|
|
expect(sections[4].headingPath).toEqual(['Users', 'Inviting users', 'Using the Auth Admin API'])
|
|
})
|
|
|
|
it('collects body text (paragraphs and lists) under each heading', () => {
|
|
expect(sections[1].content).toContain(
|
|
'Supabase distinguishes between permanent and anonymous users.'
|
|
)
|
|
expect(sections[1].content).toContain('Permanent users are tied to PII.')
|
|
expect(sections[1].content).toContain("Anonymous users aren't tied to any identities.")
|
|
})
|
|
|
|
it('does not leak content across sections', () => {
|
|
expect(sections[0].content).not.toContain('Supabase distinguishes')
|
|
expect(sections[1].content).not.toContain('invite')
|
|
})
|
|
|
|
it('includes code block text in the section body', () => {
|
|
expect(sections[4].content).toContain('const x = 1')
|
|
})
|
|
|
|
it('keeps text before the first heading as a headless section', () => {
|
|
const md = `Intro paragraph.\n\n# Title\n\nBody.`
|
|
const result = extractSections(parseMarkdownAst(md))
|
|
expect(result[0]).toEqual({
|
|
heading: '',
|
|
level: 0,
|
|
headingPath: [],
|
|
content: 'Intro paragraph.',
|
|
})
|
|
expect(result[1].headingPath).toEqual(['Title'])
|
|
})
|
|
})
|
|
|
|
describe('nodeToText', () => {
|
|
it('separates table cells and rows so words do not run together', () => {
|
|
const md = `| Attr | Type |\n| --- | --- |\n| id | string |\n| aud | string |`
|
|
const table = parseMarkdownAst(md).children[0]
|
|
expect(nodeToText(table)).toBe('Attr | Type\nid | string\naud | string')
|
|
})
|
|
|
|
it('puts list items on separate lines', () => {
|
|
const list = parseMarkdownAst('- one\n- two').children[0]
|
|
expect(nodeToText(list)).toBe('one\ntwo')
|
|
})
|
|
})
|
|
|
|
describe('parsePage', () => {
|
|
it('combines title, excerpt and sections', () => {
|
|
const page = parsePage(USERS_MD)
|
|
expect(page.title).toBe('Users')
|
|
expect(page.excerpt.startsWith('A user in Supabase Auth')).toBe(true)
|
|
expect(page.sections).toHaveLength(5)
|
|
})
|
|
})
|