Files
supabase/apps/studio/components/ui/AIAssistantPanel/DisplayBlockRenderer.tsx
Matt Rossman d143571586 feat(assistant): trace-level scorers + server-side tool execution with needsApproval (#45654)
## Motivation

When Assistant runs a potentially destructive tool like `execute_sql`,
it stops the LLM request and prompts for client-side approval and
execution of the tool. After approval, a second request kicks off under
a separate trace. This has made scoring and
[Topics](https://www.braintrust.dev/blog/topics) classification
challenging, as the generated `output` is split across stateless
requests. The [span-level
scoring](https://www.braintrust.dev/docs/evaluate/custom-code#score-spans)
approach we've used thusfar (after the LLM call, we massage the result
into an `output` payload that's stuck onto the root span) has been
cumbersome and led to invalid scores / topics where only part of the
assistant response is considered. It's also inefficient, as we're
duplicating potentially large info (like the `search_docs` output) that
already exists within the trace.

An alternative to scoring spans is to [score
traces](https://www.braintrust.dev/docs/evaluate/custom-code#score-traces).
Braintrust [best
practices](https://www.braintrust.dev/docs/evaluate/score-online#best-practices)
advise:

> Use span scope for evaluating individual operations or outputs. Use
trace scope for evaluating multi-turn conversations, overall workflow
completion, or when your scorer needs access to the full execution
context.

We've also received [direct
guidance](https://supabase.slack.com/archives/C05QYJBLX89/p1777925770927149?thread_ts=1777905716.911979&cid=C05QYJBLX89)
from their team to use this approach.

## Changes

Migrates eval scorers from custom `AssistantEvalOutput` shape to
trace-level scoring via `trace.getThread()` / `trace.getSpans()`, with
thread parsing that scores the full latest Assistant turn and passes
prior conversation separately where relevant.

Moves `execute_sql` and `deploy_edge_function` from client-side
execution after approval to AI SDK `needsApproval` + server-side
`execute()`. SQL results returned to the model are gated by AI opt-in
level, so row data is only included with `schema_and_log_and_data`;
otherwise the tool returns the no-data-permissions sentinel.

Adds `metadata.isFinalStep` to disambiguate multiple LLM requests within
an "assistant" turn due to tool call requests/responses. For online
evals, this means we should configure automations to only score traces
with `metadata.isFinalStep = true` to ensure we're judging the complete
generated response.

Other minor kaizen changes:
- Renamed `promptProviderOptions` to `systemProviderOptions` to clarify
that this is associated with the "system" message and disambiguate from
the root `providerOptions`
- Adds `evals/trace-utils.ts` to handle Zod validation of the `unknown`
span shapes from Braintrust, to more easily access typed inputs/output
on tool spans.
- Bumps AI SDK floor version `^6.0.116` → `^6.0.174`
- Tweaked the "Conciseness" scorer to not unfairly dock points for the
new `[called tool_name]` labels in serialized assistant response

## Verification

In the studio staging build, I asked Assistant to create a todos table
with 3 sample todos. I manually approved the `execute_sql` call and saw
Assistant generate text before & after the call.

In Braintrust I verified two traces were produced (see [filtered
logs](https://www.braintrust.dev/app/supabase.io/p/Assistant/logs?v=Staging&tvt=trace&search={%22filter%22:[{%22text%22:%22metadata.environment%2520%253D%2520%27staging%27%22,%22label%22:%22metadata.environment%2520%253D%2520%27staging%27%22,%22originType%22:%22btql%22},{%22text%22:%22%2560Chat%2520ID%2560%2520%253D%2520%25221cb2ac45-e5e7-458c-9da4-3bf6863b8842%2522%22,%22label%22:%22Chat%2520ID%2520equals%25201cb2ac45-e5e7-458c-9da4-3bf6863b8842%22,%22originType%22:%22form%22}]})),
the first with `metadata.isFinalStep = false` and the second with
`metadata.isFinalStep = true`.

In the Braintrust staging scorers, I ran the preview Completeness scorer
on the second trace and verified it sees the complete Assistant response
including markers for tool calls ([link to
trace](https://www.braintrust.dev/app/supabase.io/p/Assistant%20(Staging%20Scorers)/trace?object_type=project_logs&object_id=b5214b62-ad1e-4929-9d5b-40b1daebe948&r=0ed0a4f8-8aff-4a34-bb1d-1df1d88a5070&s=ff9015f8-6bf7-4ab3-83a9-ca4e69e27e82))

<img width="1193" height="960" alt="CleanShot 2026-05-07 at 11 27 10@2x"
src="https://github.com/user-attachments/assets/509d4858-c3a1-4068-986d-3aa4d5617d1a"
/>

I also tested the `deploy_edge_function` workflow and verified it still
prompts for permission and warns on deployment of existing functions.

**References**
- https://www.braintrust.dev/docs/evaluate/custom-code#score-traces
-
https://ai-sdk.dev/docs/ai-sdk-core/tools-and-tool-calling#tool-execution-approval

Supercedes https://github.com/supabase/supabase/pull/45556 and
https://github.com/supabase/supabase/pull/45339

Closes AI-473

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
* Tool actions (SQL execution, edge-function deploy) now require
explicit user Approve/Deny before proceeding.

* **Improvements**
* Assistant pauses for approval responses before sending follow-ups,
giving clearer control over risky actions.
  * Deploy/replace flows show confirmation and clearer replace warnings.
* Evaluation/scoring updated to use richer trace data for more accurate
assistant performance signals.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-05-12 15:24:21 -04:00

267 lines
8.8 KiB
TypeScript

import { acceptUntrustedSql, type UntrustedSqlFragment } from '@supabase/pg-meta'
import { PermissionAction } from '@supabase/shared-types/out/constants'
import { useQueryClient } from '@tanstack/react-query'
import type { ToolUIPart } from 'ai'
import { useParams } from 'common'
import { useRouter } from 'next/router'
import { useRef, useState, type DragEvent, type PropsWithChildren } from 'react'
import { DEFAULT_CHART_CONFIG, QueryBlock } from '../QueryBlock/QueryBlock'
import { identifyQueryType } from './AIAssistant.utils'
import { ConfirmFooter } from './ConfirmFooter'
import { ChartConfig } from '@/components/interfaces/SQLEditor/UtilityPanel/ChartConfig'
import { entityTypeKeys } from '@/data/entity-types/keys'
import { lintKeys } from '@/data/lint/keys'
import { usePrimaryDatabase } from '@/data/read-replicas/replicas-query'
import { useExecuteSqlMutation } from '@/data/sql/execute-sql-mutation'
import { useSendEventMutation } from '@/data/telemetry/send-event-mutation'
import { useChangedSync } from '@/hooks/misc/useChanged'
import { useAsyncCheckPermissions } from '@/hooks/misc/useCheckPermissions'
import { useSelectedOrganizationQuery } from '@/hooks/misc/useSelectedOrganization'
import { useProfile } from '@/lib/profile'
interface DisplayBlockRendererProps {
messageId: string
toolCallId: string
initialArgs: {
sql: UntrustedSqlFragment
label?: string
isWriteQuery?: boolean
view?: 'table' | 'chart'
xAxis?: string
yAxis?: string
}
initialResults?: unknown
/** Called when locally running SQL fails before or during client-side execution. */
onError?: (args: { messageId: string; errorText: string }) => void
/** Responds affirmatively to an AI SDK tool approval request; does not run SQL directly. */
onApprove?: () => void
/** Responds negatively to an AI SDK tool approval request; does not run SQL directly. */
onDeny?: () => void
/** AI SDK tool state used to show approval UI for pending tool calls. */
toolState?: ToolUIPart['state']
isLastPart?: boolean
isLastMessage?: boolean
showConfirmFooter?: boolean
onChartConfigChange?: (chartConfig: ChartConfig) => void
/** Called when the user clicks the query block play button to run SQL locally. */
onQueryRun?: (queryType: 'select' | 'mutation') => void
}
export const DisplayBlockRenderer = ({
messageId,
toolCallId,
initialArgs,
initialResults,
onError,
onApprove,
onDeny,
toolState,
isLastPart = false,
isLastMessage = false,
showConfirmFooter = true,
onChartConfigChange,
onQueryRun,
}: PropsWithChildren<DisplayBlockRendererProps>) => {
const queryClient = useQueryClient()
const savedInitialArgs = useRef(initialArgs)
const savedInitialResults = useRef(initialResults)
const savedInitialConfig = useRef<ChartConfig>({
...DEFAULT_CHART_CONFIG,
view: initialArgs.view === 'chart' ? 'chart' : 'table',
xKey: initialArgs.xAxis ?? '',
yKey: initialArgs.yAxis ?? '',
})
const router = useRouter()
const { ref } = useParams()
const { profile } = useProfile()
const { data: org } = useSelectedOrganizationQuery()
const { mutate: sendEvent } = useSendEventMutation()
const { can: canCreateSQLSnippet } = useAsyncCheckPermissions(
PermissionAction.CREATE,
'user_content',
{
resource: { type: 'sql', owner_id: profile?.id },
subject: { id: profile?.id },
}
)
const [chartConfig, setChartConfig] = useState<ChartConfig>(() => ({
...DEFAULT_CHART_CONFIG,
view: initialArgs.view === 'chart' ? 'chart' : 'table',
xKey: initialArgs.xAxis ?? '',
yKey: initialArgs.yAxis ?? '',
}))
const [rows, setRows] = useState<any[] | undefined>(
Array.isArray(initialResults) ? initialResults : undefined
)
const isReportsPage = router.pathname.endsWith('/reports/[id]')
const isHomePage = router.pathname === '/project/[ref]'
const isDraggableToReports = canCreateSQLSnippet && (isReportsPage || isHomePage)
const label = initialArgs.label || 'SQL Results'
const [isWriteQuery, setIsWriteQuery] = useState<boolean>(initialArgs.isWriteQuery || false)
const sqlQuery = initialArgs.sql
const { database: primaryDatabase } = usePrimaryDatabase({ projectRef: ref })
const readOnlyConnectionString = primaryDatabase?.connection_string_read_only
const postgresConnectionString = primaryDatabase?.connectionString
const {
mutate: executeSql,
error: executeSqlError,
isPending: executeSqlLoading,
} = useExecuteSqlMutation({
onError: () => {
// Suppress toast because error message is displayed inline
},
})
const toolCallIdChanged = useChangedSync(toolCallId)
if (toolCallIdChanged) {
setChartConfig(savedInitialConfig.current)
onChartConfigChange?.(savedInitialConfig.current)
setIsWriteQuery(savedInitialArgs.current.isWriteQuery || false)
setRows(Array.isArray(savedInitialResults.current) ? savedInitialResults.current : undefined)
}
const initialResultsChanged = useChangedSync(initialResults)
if (initialResultsChanged) {
const normalized = Array.isArray(initialResults) ? initialResults : undefined
if (!normalized || normalized === rows) return
setRows(normalized)
}
const handleRunQuery = (queryType: 'select' | 'mutation') => {
if (!sqlQuery) return
onQueryRun?.(queryType)
sendEvent({
action: 'assistant_suggestion_run_query_clicked',
properties: {
queryType,
...(queryType === 'mutation' ? { category: identifyQueryType(sqlQuery) ?? 'unknown' } : {}),
},
groups: {
project: ref ?? 'Unknown',
organization: org?.slug ?? 'Unknown',
},
})
}
const runQuery = (queryType: 'select' | 'mutation') => {
if (!ref || !sqlQuery) return
const connectionString =
queryType === 'mutation'
? postgresConnectionString
: (readOnlyConnectionString ?? postgresConnectionString)
if (!connectionString) {
const fallbackMessage = 'Unable to find a database connection to execute this query.'
onError?.({ messageId, errorText: fallbackMessage })
return
}
if (queryType === 'mutation') {
setIsWriteQuery(true)
}
executeSql(
{ projectRef: ref, connectionString, sql: acceptUntrustedSql(sqlQuery) },
{
onSuccess: (data) => {
setRows(Array.isArray(data.result) ? data.result : undefined)
setIsWriteQuery(queryType === 'mutation' || initialArgs.isWriteQuery || false)
if (queryType === 'mutation') {
queryClient.invalidateQueries({ queryKey: lintKeys.lint(ref) })
queryClient.invalidateQueries({ queryKey: entityTypeKeys.list(ref) })
}
},
onError: (error) => {
const lowerMessage = error.message.toLowerCase()
const isReadOnlyError =
lowerMessage.includes('read-only transaction') ||
lowerMessage.includes('permission denied') ||
lowerMessage.includes('must be owner')
if (queryType === 'select' && isReadOnlyError) {
setIsWriteQuery(true)
}
onError?.({ messageId, errorText: error.message })
},
}
)
}
const handleExecute = (queryType: 'select' | 'mutation') => {
handleRunQuery(queryType)
runQuery(queryType)
}
const handleUpdateChartConfig = ({
chartConfig: updatedValues,
}: {
chartConfig: Partial<ChartConfig>
}) => {
setChartConfig((prev) => {
const next = { ...prev, ...updatedValues }
onChartConfigChange?.(next)
return next
})
}
const handleDragStart = (e: DragEvent<Element>) => {
e.dataTransfer.setData(
'application/json',
JSON.stringify({ label, sql: sqlQuery, config: chartConfig })
)
}
const shouldShowConfirmFooter =
showConfirmFooter &&
toolState === 'approval-requested' &&
isLastPart &&
isLastMessage &&
!!onApprove &&
!!onDeny
return (
<div className="display-block w-auto overflow-x-hidden">
<div className="relative z-10">
<QueryBlock
label={label}
isWriteQuery={isWriteQuery}
sql={sqlQuery}
results={rows}
errorText={executeSqlError?.message}
chartConfig={chartConfig}
onExecute={handleExecute}
onUpdateChartConfig={handleUpdateChartConfig}
draggable={isDraggableToReports}
onDragStart={handleDragStart}
disabled={shouldShowConfirmFooter}
isExecuting={executeSqlLoading}
/>
</div>
{shouldShowConfirmFooter && (
<div className="mx-4">
<ConfirmFooter
message="Assistant wants to run this query"
cancelLabel="Skip"
confirmLabel={executeSqlLoading ? 'Running...' : 'Run Query'}
isLoading={executeSqlLoading}
onCancel={onDeny}
onConfirm={onApprove}
/>
</div>
)}
</div>
)
}