mirror of
https://github.com/supabase/supabase.git
synced 2026-10-05 09:25:06 +03:00
- Eval harness's only live tool, `search_docs`, no longer needs the in-process MCP client or its dummy token — it now calls the public docs GraphQL API (`https://supabase.com/docs/api/graphql`) directly. Low risk as this is an eval-harness change only. Production assistant path (`mcp-tools.ts`) untouched. **Update:** per [@mattrossman's review](https://github.com/supabase/supabase/pull/50092#discussion_r3980396341), the eval tool's description embeds the Content API's own GraphQL schema (fetched via a `{ schema }` query and minified with `gqlmin`), mirroring how `@supabase/mcp-server-supabase`'s `docs-tools.ts`/`loadSchema` populates production's `search_docs` description. Without it, the model had no schema to work from and issued malformed queries, which caused the 218 `search_docs` errors and the -25pp Docs Faithfulness regression in the first eval run on this PR. Schema loading is required: `createSearchDocsTool()` rejects if the schema fetch fails, so preflight and the gated eval job fail loudly instead of producing untrustworthy fallback results. `createSearchDocsTool` is async because the `ai` package's `tool()` only accepts a plain string `description`, unlike the MCP SDK's async description support; both callers (`getMockTools`, `evals/preflight.ts`) await it. `gqlmin` is a direct `apps/studio` dependency and was already transitive via `@supabase/mcp-server-supabase`. ### Verification - `pnpm -C apps/studio exec -- tsc --noEmit` reaches the compiler; it reports only the pre-existing unrelated `packages/ui-patterns/src/McpUrlBuilder/components/InstructionBlocks.tsx` `StaticImageData` error. - `pnpm -C apps/studio exec -- vitest run lib/ai/tools/mock-tools.test.ts lib/ai/tools/mcp-tools.test.ts` — 21/21 passed. - `pnpm exec tsx evals/preflight.ts` — live docs API schema fetch and search_docs call passed. - `NEXT_PUBLIC_CONTENT_API_URL=http://127.0.0.1:1/graphql pnpm -C apps/studio exec -- tsx evals/preflight.ts` — failed fast as expected, proving schema/API failures gate evals. - Fresh `run-evals` pass: Docs Faithfulness 55.7% (0pp), with no systemic `search_docs` regression. Risk: eval-harness-only; schema/API outage now fails the eval job before scoring rather than allowing fallback descriptions. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added documentation search powered by the public Supabase documentation GraphQL API. * Documentation search results now include live schema information and clearer error handling for failed or invalid requests. * **Bug Fixes** * Improved evaluation tooling reliability by removing unnecessary connection-abort behavior. * Updated validation to detect missing search tools and malformed documentation responses. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
71 lines
2.4 KiB
TypeScript
71 lines
2.4 KiB
TypeScript
import assert from 'node:assert'
|
||
import { Eval } from 'braintrust'
|
||
|
||
import { dataset } from './dataset'
|
||
import {
|
||
completenessScorer,
|
||
concisenessScorer,
|
||
correctnessScorer,
|
||
docsFaithfulnessScorer,
|
||
goalCompletionScorer,
|
||
knowledgeUsageScorer,
|
||
safetyScorer,
|
||
toolUsageScorer,
|
||
urlValidityScorer,
|
||
} from './scorer'
|
||
import { sqlIdentifierQuotingScorer, sqlSyntaxScorer } from './scorer-wasm'
|
||
import { buildTranscript } from './transcript'
|
||
import { generateAssistantResponse } from '@/lib/ai/generate-assistant-response'
|
||
import { getModel } from '@/lib/ai/model'
|
||
import { DEFAULT_ASSISTANT_BASE_MODEL_ID, getAssistantModelEntry } from '@/lib/ai/model.utils'
|
||
import { getMockTools } from '@/lib/ai/tools/mock-tools'
|
||
|
||
assert(process.env.BRAINTRUST_PROJECT_ID, 'BRAINTRUST_PROJECT_ID is not set')
|
||
assert(process.env.OPENAI_API_KEY, 'OPENAI_API_KEY is not set')
|
||
|
||
Eval('Assistant', {
|
||
projectId: process.env.BRAINTRUST_PROJECT_ID,
|
||
trialCount: process.env.CI ? 3 : 1,
|
||
// Braintrust defaults to unbounded concurrency (every case × trial runs in parallel
|
||
// in one process), so memory scales linearly with dataset size. Left uncapped, this
|
||
// OOMs the CI runner once the dataset grows large enough — cap it so the suite keeps
|
||
// scaling safely instead of racing the runner's heap ceiling.
|
||
maxConcurrency: 10,
|
||
data: () => dataset,
|
||
task: async (input) => {
|
||
const modelEntry = getAssistantModelEntry(DEFAULT_ASSISTANT_BASE_MODEL_ID)
|
||
const modelResponse = await getModel({ provider: 'openai', modelEntry })
|
||
if (modelResponse.error) throw modelResponse.error
|
||
|
||
const result = await generateAssistantResponse({
|
||
...modelResponse.modelParams,
|
||
isExplorerEnabled: true,
|
||
messages: [
|
||
{
|
||
id: '1',
|
||
role: 'user',
|
||
parts: [{ type: 'text', text: input.prompt }],
|
||
},
|
||
],
|
||
tools: await getMockTools(input.mockTables ? { list_tables: input.mockTables } : undefined),
|
||
})
|
||
|
||
const finishReason = await result.finishReason
|
||
const steps = await result.steps
|
||
return { finishReason, transcript: buildTranscript(input.prompt, steps) }
|
||
},
|
||
scores: [
|
||
toolUsageScorer,
|
||
knowledgeUsageScorer,
|
||
sqlSyntaxScorer,
|
||
sqlIdentifierQuotingScorer,
|
||
goalCompletionScorer,
|
||
concisenessScorer,
|
||
completenessScorer,
|
||
docsFaithfulnessScorer,
|
||
correctnessScorer,
|
||
safetyScorer,
|
||
urlValidityScorer,
|
||
],
|
||
})
|