Files
Pedro RodriguesandClaude Opus 4.8 47595f8ac7 feat(self-hosted): implement queryLogs for the MCP debugging tools (#48900)
> [!IMPORTANT]  
>
> Only merge this when (https://github.com/supabase/platform/pull/36804)
is merged, as the AI assistant will not have access to the `query_logs`
tool for the remote MCP server

## I have read the
[CONTRIBUTING.md](https://github.com/supabase/supabase/blob/master/CONTRIBUTING.md)
file.

YES

## What kind of change does this PR introduce?

Feature (self-hosted / CLI Studio MCP server).

## What is the current behavior?

Self-hosted `getDebuggingOperations`
(`apps/studio/lib/api/self-hosted/mcp.ts`) implements only `getLogs`, so
the MCP `debugging` group exposes `get_logs` — a fixed per-service log
dump built by `getLogQuery`. Logs are served by Logflare, which speaks
BigQuery SQL.

## What is the new behavior?

Bumps `@supabase/mcp-server-supabase` to `^0.10.0` (adds `query_logs` +
`logsDialect`, and hides `get_logs` wherever a platform declares
`queryLogs`) and moves logs over to it.

- **Self-hosted `query_logs`:** declares `logsDialect: 'bigquery'` and
implements `queryLogs`, passing the model's SQL straight through to the
same Logflare `logs.all` endpoint (arbitrary `sql` param) — no new
endpoint, no dialect translation.
- **Drops `get_logs` from self-hosted:** `getLogs` throws (the server
hides it once `queryLogs` exists) and the per-service `getLogQuery`
builder is deleted; the model now writes its own BigQuery SQL, guided by
the dialect schema hint.
- **Honors no-logs mode:** `query_logs` throws when `logs:all` is
disabled — the self-hosted default, enabled via the
`docker-compose.logs.yml` override.
- **Assistant:** switches the dashboard assistant from `get_logs` to
`query_logs` (allowlist, drift guard, prompt, mocks, evals).

Refs AI-1046


<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->
## Summary by CodeRabbit

* **New Features**
  * AI debugging can query recent project logs using read-only SQL.
* Log queries support optional time-range filters, filtering,
aggregation, and joins.
* Self-hosted debugging checks whether logging is enabled before running
queries.

* **Bug Fixes**
* Updated debugging workflows and validation to consistently use the new
log-query capability.
* Removed reliance on legacy service-specific log filtering and query
behavior.

* **Documentation**
* Updated MCP debugging tool guidance to describe SQL-based log queries.

* **Tests**
* Expanded coverage for enabled, disabled, and unsupported logging
scenarios.
<!-- end of auto-generated comment: release notes by coderabbit.ai -->

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
2026-08-12 13:11:23 +01:00
..
2025-12-22 23:45:48 -05:00

Studio Assistant Evals

We use Braintrust to evaluate Assistant behaviors against a tracked dataset (offline evals) and against live traces (online evals).

Offline Evals

Add offline eval test cases to dataset.ts. If needed, add new scorers (see below) for the specific dimension you wish to test. Expect to update and run offline evals when adding new Assistant behaviors

You may wish to run offline evals when:

  • You updated the eval suite with a new test case or scorer
  • You changed Assistant's behavior and want to check for improvements/regressions

Running Offline Evals in CI

Add the run-evals label on a PR to the repo and Braintrust's GitHub Action will run evals and post a summary comment (example).

You can find detailed results in the "Experiments" tab of the "Assistant" project on Braintrust.

Running Offline Evals in Local Dev

Within apps/studio

# To set up WASM files
pnpm evals:setup

# Run all evals and upload results to Braintrust
pnpm evals:upload

# Run all evals without uploading results
pnpm evals:run

# Run an upload single test case
pnpm braintrust eval evals/assistant.eval.ts --filter "input.prompt=How many projects"

Upload results when you want to inspect Experiments or Logs in the Braintrust dashboard or API. You can use developer tools like Braintrust MCP or bt CLI to analyze results with an agent.

Scorers

Scorers look at a thread or task output and assign a score deterministically or via LLM-as-a-judge. Optionally they can consider expected values.

Define scorers in scorer.ts and include them in assistant.eval.ts to run them in offline evals.

Updating Online Scorers

Online scorers run as serverless functions on Braintrust infrastructure. They're deployed from the scorer-online.ts script. Since these scoring against production traces, they can't rely on ground truth expected values. Structure scoring logic and LLM prompts accordingly. Not every scorer needs to be an online scorer.

To opt-in to online scoring, add the scorer to scorer-online-manifest.json and add a corresponding handler in scorer-online.ts

Testing & Deploying Online Scorers

Add the preview-scorers label to a PR to deploy branch-prefixed scorers to the "Assistant (Staging Scorers)" Braintrust project (example). From that project dashboard, you can manually test the scorer against a trace from any project.

After merge to master, preview scorers automatically clean up and deploy to the production in the "Assistant" Braintrust project. Update the "Online Scoring" automation in the Logs page to include the new scorer function.