Files
supabase/apps/docs/scripts/search_v2/ingest.ts
T
Jeremias Menichelli 7f1df457f8 feat: Create action to upsert table for new search (#50683)
## Problem

For the new search, we will scan all files present within the
public/markdown directory, extract text nodes and upsert a new Supabase
table in a new project to do FTS type search.

## Solution

In this PR:
- A new script directory is created with files to solve all the steps
described above.
 - Unit tests added for the fundamental bits of the script.
- A new workflow file is added so the action runs after push every time
content is altered, added or removed.

<!--
## Preview links

If relevant, include links to changed pages for easy review access.

Copy the preview base URL from the Vercel bot comment on this PR. Use
the following table as an example template.

| Site | Live | Preview | Search for |
| -------------- |
-------------------------------------------------------------------------
|
------------------------------------------------------------------------------------------------------------
| ----------------------------- |
| WWW | [/blog/your-post](https://supabase.com/blog/your-post) |
[/blog/your-post](https://zone-www-dot-com-git-branch-name-supabase.vercel.app/blog/your-post)
| unique phrase from the change |
| Docs |
[/docs/guides/your-page](https://supabase.com/docs/guides/your-page) |
[/docs/guides/your-page](https://docs-git-branch-name-supabase.vercel.app/docs/guides/your-page)
| unique phrase from the change |
| Studio | [/dashboard](https://supabase.com/dashboard) |
[/dashboard](https://studio-git-branch-name-supabase.vercel.app/dashboard)
| unique phrase from the change |
| Design system | [/design-system](https://supabase.com/design-system) |
[/design-system](https://design-system-git-branch-name-supabase.vercel.app/design-system)
| unique phrase from the change |
| UI library | [/library](https://supabase.com/library) |
[/library](https://ui-library-git-branch-name-supabase.vercel.app/library)
| unique phrase from the change |
| Knowledge base |
[/kb/guides/your-page](https://supabase.com/kb/guides/your-page) |
[/kb/guides/your-page](https://kb-git-branch-name-supabase.vercel.app/kb/guides/your-page)
| unique phrase from the change |
-->

<!-- ## Additional context

Optionally add any other context or screenshots.

-->

## Review instructions

Sadly, testing this work is quite complex, but in case someone wants to:

1. Create a new Supabase project ton your personal space
1. Copy the id of the project and a secret key and add it to the new
Search V2 environment variables as shown in the example file
1. Copy the content of the newly added `setup.sql` and run it on the SQL
editor of your project.
1. Fetch this branch and on the search directory, run `search-v2:ingest`
1. Your project's table should have rows corresponding to the content
from docs


## Checklist

Check all before review:

- [x] I have read
[CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md)
- [x] If I wrote a new docs topic or edited an existing topic, I used
the `/write-the-docs` or `/edit-the-docs` skill, which references
[WORD_LIST](https://github.com/supabase/supabase/blob/master/apps/docs/WORD_LIST.md)
and the docs
[CONTRIBUTING](https://github.com/supabase/supabase/blob/master/apps/docs/CONTRIBUTING.md)
guide


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Documentation pages are now indexed for full-text search at the page
and section level.
- Search results can show the most relevant section from each page, with
its title, heading, excerpt, and relevance score.
- Search content is automatically refreshed when published Markdown
documentation changes.
- **Tests**
- Added coverage for Markdown parsing, page structure, routing, and
search-content generation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-09-28 16:34:46 -03:00

113 lines
3.7 KiB
TypeScript

import '../utils/dotenv.js'
import { readFile } from 'node:fs/promises'
import path from 'node:path'
import { fileURLToPath } from 'node:url'
import fg from 'fast-glob'
import { parsePage } from './markdown.js'
import { filePathToSlug } from './routes.js'
import { createSupabaseClient, TABLE_NAME } from './supabase.js'
// ---------------------------------------------------------------------------
// Configuration
// ---------------------------------------------------------------------------
/** Folder that holds the markdown files. Searched recursively. */
export const CONTENT_ROOT = path.resolve(process.cwd(), 'public/markdown')
/** Glob patterns (relative to CONTENT_ROOT) for the files to ingest. */
const FILE_PATTERNS = ['**/*.md', '**/*.mdx']
/** Glob patterns (relative to CONTENT_ROOT) for files/folders to skip. */
const IGNORED_PATTERNS = ['reference/**', 'guides/troubleshooting/**', '_partials/**']
/** How many rows to send to Supabase per insert request. */
const INSERT_BATCH_SIZE = 200
// ---------------------------------------------------------------------------
interface SectionRow {
slug: string
file_path: string
page_title: string
heading: string
heading_level: number
heading_path: string[]
content: string
excerpt: string
}
/** Build the DB rows for a single markdown file. Pure: no I/O. */
export function buildRows(markdown: string, filePath: string, contentRoot: string): SectionRow[] {
const slug = filePathToSlug(filePath, contentRoot)
const relativePath = path.relative(contentRoot, filePath).split(path.sep).join('/')
const page = parsePage(markdown)
const pageTitle = page.title || path.basename(slug) || 'Untitled'
return page.sections.map((section) => ({
slug,
file_path: relativePath,
page_title: pageTitle,
heading: section.heading,
heading_level: section.level,
heading_path: section.headingPath,
content: section.content,
excerpt: page.excerpt,
}))
}
function chunk<T>(items: T[], size: number): T[][] {
const out: T[][] = []
for (let i = 0; i < items.length; i += size) out.push(items.slice(i, i + size))
return out
}
async function main() {
const supabase = createSupabaseClient()
const files = await fg(FILE_PATTERNS, {
cwd: CONTENT_ROOT,
absolute: true,
onlyFiles: true,
ignore: IGNORED_PATTERNS,
})
if (files.length === 0) {
console.warn(`No markdown files found under ${CONTENT_ROOT}`)
return
}
console.log(`Found ${files.length} markdown file(s) under ${CONTENT_ROOT}`)
let totalRows = 0
for (const file of files.sort()) {
const markdown = await readFile(file, 'utf8')
const rows = buildRows(markdown, file, CONTENT_ROOT)
const slug = rows[0]?.slug ?? filePathToSlug(file, CONTENT_ROOT)
// Replace everything previously ingested for this page.
const { error: deleteError } = await supabase.from(TABLE_NAME).delete().eq('slug', slug)
if (deleteError) throw new Error(`Failed to clear "${slug}": ${deleteError.message}`)
for (const batch of chunk(rows, INSERT_BATCH_SIZE)) {
const { error: insertError } = await supabase.from(TABLE_NAME).insert(batch)
if (insertError) throw new Error(`Failed to insert "${slug}": ${insertError.message}`)
}
totalRows += rows.length
console.log(` ✓ ${slug || '(root)'} (${rows.length} section${rows.length === 1 ? '' : 's'})`)
}
console.log(`Done. Ingested ${totalRows} section(s) from ${files.length} file(s).`)
}
// Only run when executed directly (so buildRows can be imported by tests).
const isDirectRun =
!!process.argv[1] && path.resolve(process.argv[1]) === fileURLToPath(import.meta.url)
if (isDirectRun) {
main().catch((err) => {
console.error(err instanceof Error ? err.message : err)
process.exit(1)
})
}