Files
supabase/apps/studio/evals/scorer-online.ts
Charis 8c409e2df5 Fix eval scorer truncation via local transcript capture (#49151)
## Problem

Scorers previously derived the assistant's final answer via Braintrust's
`trace.getThread()`, which silently truncates long traces at the
backend's preview-length cap (~10KB). The SDK never passes
`preview_length` in its BTQL query and there's no supported override.
This caused false-negative scores (Completeness, Correctness, Goal
Completion, Safety collapsing to 0/null) specifically on multi-step
tool-calling eval cases, since longer traces are more likely to have
their tail (the final assistant message) truncated away.

## Solution

Capture the assistant's full, untruncated final answer directly in the
eval task's output in memory (via AI SDK's `result.steps`, already fully
available once the stream is consumed) instead of round-tripping through
Braintrust's truncating storage/query layer. Scorers now read
`output.transcript` instead of calling `trace.getThread()`.

## Changes

- **New**: `apps/studio/evals/transcript.ts` — `Transcript` type and
`buildTranscript()` function
- **New**: `apps/studio/evals/transcript.test.ts` — unit tests (5
passing)
- **Modified**: `apps/studio/evals/assistant.eval.ts` — captures
`result.steps` and returns transcript
- **Modified**: `apps/studio/evals/scorer.ts` — migrated 7 scorers to
read from local transcript
- **Modified**: `apps/studio/evals/trace-utils.ts` — removed dead
thread-serialization code
- **Deleted**: `apps/studio/evals/trace-utils.test.ts` — superseded by
transcript tests

## Test Plan

- [x] `pnpm --filter studio typecheck` — clean
- [x] `pnpm --filter studio lint` — clean  
- [x] `npx vitest run evals/transcript.test.ts` — 5/5 passing
- [x] Full live eval run (35/35 cases) against Braintrust —
[experiment](https://www.braintrust.dev/app/supabase.io/p/Assistant/experiments/eval-scorer-transcript-capture-1786985352)
shows Completeness/Correctness/Goal Completion/Safety scores comparable
to baseline

## Known Residual Risk

Other scorers that derive data from `trace.getSpans()` (toolUsageScorer,
sqlSyntaxScorer, sqlIdentifierQuotingScorer, knowledgeUsageScorer, and
docsFaithfulnessScorer's docs-content lookup) could theoretically hit
the same truncation issue, but have not been observed to fail in
practice. This is not addressed in this PR.

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Added transcript generation from assistant interaction steps,
including text and tool-call inputs.
* Evaluation results can now include complete transcripts for detailed
conversation analysis.
* Online evaluations can derive transcripts from recorded interaction
traces when needed.

* **Bug Fixes**
* Improved scoring by selecting the appropriate conversation content for
each evaluation.
* Ensured offline transcripts take precedence when available, with
trace-based fallback support.

* **Tests**
* Added coverage for multi-step interactions, tool calls, filtering,
empty steps, and URL validation.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-08-18 12:22:50 -04:00

63 lines
2.3 KiB
TypeScript

/**
* Entry point for `braintrust push` to deploy scorers to Braintrust.
*
* Excluded scorers:
* - sqlSyntaxScorer, sqlIdentifierQuotingScorer: use libpg-query (WASM),
* which esbuild cannot bundle for Braintrust's remote infra.
* - toolUsageScorer: requires expected.requiredTools, offline-eval-only.
* - correctnessScorer: requires ground truth (expected output), offline-eval-only.
*/
import braintrust from 'braintrust'
import {
completenessScorer,
concisenessScorer,
docsFaithfulnessScorer,
goalCompletionScorer,
safetyScorer,
urlValidityScorer,
type AssistantEvalScorer,
} from './scorer'
import manifest from './scorer-online-manifest.json'
const projectId = process.env.BRAINTRUST_PROJECT_ID
if (!projectId && process.env.IS_BRAINTRUST_PUSH)
throw new Error('BRAINTRUST_PROJECT_ID is not set')
// When running in CI, prefix scorers with the branch name to avoid collisions between PRs
// in the staging project. GITHUB_HEAD_REF is only set on PR events, not push events (e.g. master).
const branch = process.env.GITHUB_HEAD_REF || process.env.GITHUB_REF_NAME
const prNumber = process.env.GITHUB_PR_NUMBER ? Number(process.env.GITHUB_PR_NUMBER) : undefined
const prefix = process.env.GITHUB_HEAD_REF
? `${process.env.GITHUB_HEAD_REF.replace(/[^a-z0-9-]/gi, '-').toLowerCase()}-`
: ''
const metadata = branch ? { gitBranch: branch, ...(prNumber && { prNumber }) } : undefined
const description = prNumber && branch ? `#${prNumber} · ${branch}` : branch
const handlers = {
'goal-completion': goalCompletionScorer,
conciseness: concisenessScorer,
completeness: completenessScorer,
'docs-faithfulness': docsFaithfulnessScorer,
'url-validity': urlValidityScorer,
// safetyScorer guards on requiresSafetyCheck (an offline-eval concept).
// Online traces have no expected, so we wrap it to always run.
safety: (args) =>
safetyScorer({ ...args, expected: { ...args.expected, requiresSafetyCheck: true } }),
} satisfies Record<string, AssistantEvalScorer>
// @ts-expect-error - Project ID is only required at build-time
const project = braintrust.projects.create({ id: projectId })
for (const { slug, name } of manifest) {
project.scorers.create({
slug: `${prefix}${slug}`,
name,
description,
handler: handlers[slug as keyof typeof handlers],
ifExists: 'replace',
metadata,
})
}