Files
supabase/apps/studio/evals/assistant.eval.ts
Pedro RodriguesandClaude Sonnet 5 22d7bc0cfd feat(studio-evals): custom search_docs tool for the eval harness (no token / no PAT) (#50092)
- Eval harness's only live tool, `search_docs`, no longer needs the
in-process MCP client or its dummy token — it now calls the public docs
GraphQL API (`https://supabase.com/docs/api/graphql`) directly. Low risk
as this is an eval-harness change only. Production assistant path
(`mcp-tools.ts`) untouched.

**Update:** per [@mattrossman's
review](https://github.com/supabase/supabase/pull/50092#discussion_r3980396341),
the eval tool's description embeds the Content API's own GraphQL schema
(fetched via a `{ schema }` query and minified with `gqlmin`), mirroring
how `@supabase/mcp-server-supabase`'s `docs-tools.ts`/`loadSchema`
populates production's `search_docs` description. Without it, the model
had no schema to work from and issued malformed queries, which caused
the 218 `search_docs` errors and the -25pp Docs Faithfulness regression
in the first eval run on this PR. Schema loading is required:
`createSearchDocsTool()` rejects if the schema fetch fails, so preflight
and the gated eval job fail loudly instead of producing untrustworthy
fallback results. `createSearchDocsTool` is async because the `ai`
package's `tool()` only accepts a plain string `description`, unlike the
MCP SDK's async description support; both callers (`getMockTools`,
`evals/preflight.ts`) await it. `gqlmin` is a direct `apps/studio`
dependency and was already transitive via
`@supabase/mcp-server-supabase`.

### Verification
- `pnpm -C apps/studio exec -- tsc --noEmit` reaches the compiler; it
reports only the pre-existing unrelated
`packages/ui-patterns/src/McpUrlBuilder/components/InstructionBlocks.tsx`
`StaticImageData` error.
- `pnpm -C apps/studio exec -- vitest run
lib/ai/tools/mock-tools.test.ts lib/ai/tools/mcp-tools.test.ts` — 21/21
passed.
- `pnpm exec tsx evals/preflight.ts` — live docs API schema fetch and
search_docs call passed.
- `NEXT_PUBLIC_CONTENT_API_URL=http://127.0.0.1:1/graphql pnpm -C
apps/studio exec -- tsx evals/preflight.ts` — failed fast as expected,
proving schema/API failures gate evals.
- Fresh `run-evals` pass: Docs Faithfulness 55.7% (0pp), with no
systemic `search_docs` regression.

Risk: eval-harness-only; schema/API outage now fails the eval job before
scoring rather than allowing fallback descriptions.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **New Features**
* Added documentation search powered by the public Supabase
documentation GraphQL API.
* Documentation search results now include live schema information and
clearer error handling for failed or invalid requests.

* **Bug Fixes**
* Improved evaluation tooling reliability by removing unnecessary
connection-abort behavior.
* Updated validation to detect missing search tools and malformed
documentation responses.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Sonnet 5 <noreply@anthropic.com>
2026-09-14 12:51:14 +02:00

71 lines
2.4 KiB
TypeScript
Raw Permalink Blame History

This file contains ambiguous Unicode characters
This file contains Unicode characters that might be confused with other characters. If you think that this is intentional, you can safely ignore this warning. Use the Escape button to reveal them.
import assert from 'node:assert'
import { Eval } from 'braintrust'
import { dataset } from './dataset'
import {
completenessScorer,
concisenessScorer,
correctnessScorer,
docsFaithfulnessScorer,
goalCompletionScorer,
knowledgeUsageScorer,
safetyScorer,
toolUsageScorer,
urlValidityScorer,
} from './scorer'
import { sqlIdentifierQuotingScorer, sqlSyntaxScorer } from './scorer-wasm'
import { buildTranscript } from './transcript'
import { generateAssistantResponse } from '@/lib/ai/generate-assistant-response'
import { getModel } from '@/lib/ai/model'
import { DEFAULT_ASSISTANT_BASE_MODEL_ID, getAssistantModelEntry } from '@/lib/ai/model.utils'
import { getMockTools } from '@/lib/ai/tools/mock-tools'
assert(process.env.BRAINTRUST_PROJECT_ID, 'BRAINTRUST_PROJECT_ID is not set')
assert(process.env.OPENAI_API_KEY, 'OPENAI_API_KEY is not set')
Eval('Assistant', {
projectId: process.env.BRAINTRUST_PROJECT_ID,
trialCount: process.env.CI ? 3 : 1,
// Braintrust defaults to unbounded concurrency (every case × trial runs in parallel
// in one process), so memory scales linearly with dataset size. Left uncapped, this
// OOMs the CI runner once the dataset grows large enough — cap it so the suite keeps
// scaling safely instead of racing the runner's heap ceiling.
maxConcurrency: 10,
data: () => dataset,
task: async (input) => {
const modelEntry = getAssistantModelEntry(DEFAULT_ASSISTANT_BASE_MODEL_ID)
const modelResponse = await getModel({ provider: 'openai', modelEntry })
if (modelResponse.error) throw modelResponse.error
const result = await generateAssistantResponse({
...modelResponse.modelParams,
isExplorerEnabled: true,
messages: [
{
id: '1',
role: 'user',
parts: [{ type: 'text', text: input.prompt }],
},
],
tools: await getMockTools(input.mockTables ? { list_tables: input.mockTables } : undefined),
})
const finishReason = await result.finishReason
const steps = await result.steps
return { finishReason, transcript: buildTranscript(input.prompt, steps) }
},
scores: [
toolUsageScorer,
knowledgeUsageScorer,
sqlSyntaxScorer,
sqlIdentifierQuotingScorer,
goalCompletionScorer,
concisenessScorer,
completenessScorer,
docsFaithfulnessScorer,
correctnessScorer,
safetyScorer,
urlValidityScorer,
],
})