Files
supabase/apps/studio/evals
Saxon Fletcher aa2897f712 feat(studio): teach assistant to query ClickHouse logs (#49292)
## I have read the
[CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md)
file.

YES

## What kind of change does this PR introduce?

Feature and bug fix.

## What is the current behavior?

The assistant can call `query_logs`, but it is not given the ClickHouse
schema and query-writing guidance it needs. It also lacks a current UTC
reference for producing the absolute timestamps required by the tool,
which can lead to valid queries being run against the wrong time range
and reported as returning zero rows.

## What is the new behavior?

- Adds a dedicated `logs` knowledge topic backed by the shared
ClickHouse schema and query guidance.
- Requires the assistant to load that knowledge before using
`query_logs`.
- Includes the current UTC time in project context so relative requests
can be converted to correct absolute tool parameters.
- Covers the new knowledge flow and context with focused tests and
updates the assistant eval expectation.

## How to test

1. Check out this PR and run Studio against a project that has recent
logs. Generate some project activity first, such as an API request, if
needed.
2. Open the AI Assistant and ask: `Show log counts by minute for the
last 15 minutes and summarize any spikes.`
3. Expand the assistant's tool activity and verify it loads the `logs`
knowledge topic before calling `query_logs`.
4. Inspect the `query_logs` input and verify:
- `iso_timestamp_start` and `iso_timestamp_end` are absolute UTC
timestamps ending in `Z`.
   - The timestamps cover approximately the requested 15-minute window.
- The SQL uses ClickHouse syntax, includes a `LIMIT`, and does not put
the time range in the SQL `WHERE` clause.
5. Verify the assistant's summary reflects the rows returned by
`query_logs` instead of reporting zero rows when results are present.

## Additional context

This is the bottom PR in stack #49294. The front-end visualization is
added separately in #49293.

Verified with 59 focused tests across assistant context, Studio/MCP
tools, query display, and logs result parsing.


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

- **New Features**
- Added AI-assisted project log querying through the `query_logs` tool.
- Added logs knowledge guidance for time ranges, schema discovery, query
limits, and concise result summaries.
- Project context now includes the current UTC timestamp to improve
relative time-range interpretation.
- Improved notebook assistance with safer table verification and
appropriate handling of log queries.

- **Bug Fixes**
- Prevented incorrect SQL timestamp filtering and enabled cross-service
searches without requiring a source filter.
  - Added validation for supported knowledge topics.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->
2026-08-21 09:30:54 +10:00
..
2025-12-22 23:45:48 -05:00

Studio Assistant Evals

We use Braintrust to evaluate Assistant behaviors against a tracked dataset (offline evals) and against live traces (online evals).

Offline Evals

Add offline eval test cases to dataset.ts. If needed, add new scorers (see below) for the specific dimension you wish to test. Expect to update and run offline evals when adding new Assistant behaviors

You may wish to run offline evals when:

  • You updated the eval suite with a new test case or scorer
  • You changed Assistant's behavior and want to check for improvements/regressions

Running Offline Evals in CI

Add the run-evals label on a PR to the repo and Braintrust's GitHub Action will run evals and post a summary comment (example).

You can find detailed results in the "Experiments" tab of the "Assistant" project on Braintrust.

Running Offline Evals in Local Dev

Within apps/studio

# To set up WASM files
pnpm evals:setup

# Run all evals and upload results to Braintrust
pnpm evals:upload

# Run all evals without uploading results
pnpm evals:run

# Run an upload single test case
pnpm braintrust eval evals/assistant.eval.ts --filter "input.prompt=How many projects"

Upload results when you want to inspect Experiments or Logs in the Braintrust dashboard or API. You can use developer tools like Braintrust MCP or bt CLI to analyze results with an agent.

Scorers

Scorers look at a thread or task output and assign a score deterministically or via LLM-as-a-judge. Optionally they can consider expected values.

Define scorers in scorer.ts and include them in assistant.eval.ts to run them in offline evals.

Updating Online Scorers

Online scorers run as serverless functions on Braintrust infrastructure. They're deployed from the scorer-online.ts script. Since these scoring against production traces, they can't rely on ground truth expected values. Structure scoring logic and LLM prompts accordingly. Not every scorer needs to be an online scorer.

To opt-in to online scoring, add the scorer to scorer-online-manifest.json and add a corresponding handler in scorer-online.ts

Testing & Deploying Online Scorers

Add the preview-scorers label to a PR to deploy branch-prefixed scorers to the "Assistant (Staging Scorers)" Braintrust project (example). From that project dashboard, you can manually test the scorer against a trace from any project.

After merge to master, preview scorers automatically clean up and deploy to the production in the "Assistant" Braintrust project. Update the "Online Scoring" automation in the Logs page to include the new scorer function.