mirror of
https://github.com/supabase/supabase.git
synced 2026-10-05 17:35:10 +03:00
- Eval harness's only live tool, `search_docs`, no longer needs the in-process MCP client or its dummy token — it now calls the public docs GraphQL API (`https://supabase.com/docs/api/graphql`) directly. Low risk as this is an eval-harness change only. Production assistant path (`mcp-tools.ts`) untouched. **Update:** per [@mattrossman's review](https://github.com/supabase/supabase/pull/50092#discussion_r3980396341), the eval tool's description embeds the Content API's own GraphQL schema (fetched via a `{ schema }` query and minified with `gqlmin`), mirroring how `@supabase/mcp-server-supabase`'s `docs-tools.ts`/`loadSchema` populates production's `search_docs` description. Without it, the model had no schema to work from and issued malformed queries, which caused the 218 `search_docs` errors and the -25pp Docs Faithfulness regression in the first eval run on this PR. Schema loading is required: `createSearchDocsTool()` rejects if the schema fetch fails, so preflight and the gated eval job fail loudly instead of producing untrustworthy fallback results. `createSearchDocsTool` is async because the `ai` package's `tool()` only accepts a plain string `description`, unlike the MCP SDK's async description support; both callers (`getMockTools`, `evals/preflight.ts`) await it. `gqlmin` is a direct `apps/studio` dependency and was already transitive via `@supabase/mcp-server-supabase`. ### Verification - `pnpm -C apps/studio exec -- tsc --noEmit` reaches the compiler; it reports only the pre-existing unrelated `packages/ui-patterns/src/McpUrlBuilder/components/InstructionBlocks.tsx` `StaticImageData` error. - `pnpm -C apps/studio exec -- vitest run lib/ai/tools/mock-tools.test.ts lib/ai/tools/mcp-tools.test.ts` — 21/21 passed. - `pnpm exec tsx evals/preflight.ts` — live docs API schema fetch and search_docs call passed. - `NEXT_PUBLIC_CONTENT_API_URL=http://127.0.0.1:1/graphql pnpm -C apps/studio exec -- tsx evals/preflight.ts` — failed fast as expected, proving schema/API failures gate evals. - Fresh `run-evals` pass: Docs Faithfulness 55.7% (0pp), with no systemic `search_docs` regression. Risk: eval-harness-only; schema/API outage now fails the eval job before scoring rather than allowing fallback descriptions. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added documentation search powered by the public Supabase documentation GraphQL API. * Documentation search results now include live schema information and clearer error handling for failed or invalid requests. * **Bug Fixes** * Improved evaluation tooling reliability by removing unnecessary connection-abort behavior. * Updated validation to detect missing search tools and malformed documentation responses. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
60 lines
2.4 KiB
TypeScript
60 lines
2.4 KiB
TypeScript
/**
|
|
* Eval preflight — search_docs connectivity check.
|
|
*
|
|
* The assistant eval harness (`getMockTools`) mocks every tool except
|
|
* `search_docs`, which is a self-contained tool that calls the public Supabase
|
|
* docs GraphQL API directly (no MCP server, no access token). If that call is
|
|
* broken (endpoint down, contract drift, missing tool), evals fail deep inside
|
|
* a Braintrust run with an opaque per-case error.
|
|
*
|
|
* This preflight exercises the exact same tool and fails fast with an
|
|
* actionable message, so a broken docs API connection is caught up front when
|
|
* the eval job runs (e.g. on push). Keep it in lockstep with how `getMockTools`
|
|
* obtains `search_docs` — both use `createSearchDocsTool` from
|
|
* `lib/ai/tools/search-docs-tool`.
|
|
*/
|
|
import { createSearchDocsTool } from '@/lib/ai/tools/search-docs-tool'
|
|
|
|
async function runPreflight() {
|
|
const searchDocs = await createSearchDocsTool()
|
|
|
|
if (!searchDocs?.execute) {
|
|
throw new Error(
|
|
'`search_docs` is missing from the eval harness. The tool contract may have ' +
|
|
'drifted, or `createSearchDocsTool` was removed from lib/ai/tools/search-docs-tool.'
|
|
)
|
|
}
|
|
|
|
const output = (await searchDocs.execute(
|
|
{
|
|
graphql_query:
|
|
'{ searchDocs(query: "row level security", limit: 1) { nodes { title href } } }',
|
|
},
|
|
{ toolCallId: 'preflight', messages: [], context: {} }
|
|
)) as { content: Array<{ type?: 'text'; text: string }> }
|
|
|
|
// Validate the MCP text-content shape the scorers parse
|
|
// (mcpTextContentSpanOutputSchema / docsFaithfulnessScorer).
|
|
const content = output?.content
|
|
const text = content?.[0]?.text
|
|
if (!Array.isArray(content) || content[0]?.type !== 'text' || typeof text !== 'string' || !text) {
|
|
throw new Error(
|
|
'`search_docs` returned an unexpected shape. Expected MCP text content ' +
|
|
'({ content: [{ type: "text", text: string }] }) but got: ' +
|
|
JSON.stringify(output ?? null)
|
|
)
|
|
}
|
|
|
|
console.log('✅ Eval preflight OK — `search_docs` reaches the public docs GraphQL API.')
|
|
}
|
|
|
|
runPreflight().catch((error) => {
|
|
console.error(
|
|
'❌ Eval preflight failed — `search_docs` cannot reach the public docs GraphQL API, ' +
|
|
'so evals would fail. Check NEXT_PUBLIC_CONTENT_API_URL (if set) and the default ' +
|
|
'endpoint https://supabase.com/docs/api/graphql.'
|
|
)
|
|
console.error(error instanceof Error ? error.message : error)
|
|
process.exit(1)
|
|
})
|