mirror of
https://github.com/supabase/supabase.git
synced 2026-10-05 17:35:10 +03:00
docs(skills): add clickhouse-logs-queries skill (#47388)
## What Adds an agent skill, `clickhouse-logs-queries`, to help teammates write and migrate logs queries against the ClickHouse-backed `logs` table. It covers: - The `logs` table schema, sources, and the `log_attributes` map - ClickHouse vs BigQuery functions (`count()`, `match`/`ilike`, `toInt32OrZero`, `mapKeys`) - Best practices (filter by source, always LIMIT, tight time range) - A BigQuery-to-ClickHouse migration guide with a full before/after - How to wire branded analytics SQL in the Studio codebase (`safeSql`/`analyticsLiteral`, the endpoint/builder pickers, the OTEL generators) ## Why The logs backend is moving to a single ClickHouse table behind the `otelLegacyLogs` flag. This skill gives a single, accurate reference so query authoring and code migration stay consistent. ## Files - `.claude/skills/clickhouse-logs-queries/SKILL.md` - `.claude/skills/clickhouse-logs-queries/references/bigquery-migration.md` - `.claude/skills/clickhouse-logs-queries/references/codebase-integration.md` Docs only, no runtime code. <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **New Features** * Added guidance for working with ClickHouse-backed logs queries, including source filtering, structured field access, function equivalents, and example queries. * Added a step-by-step reference for converting existing logs SQL to the new query format. * **Documentation** * Added implementation notes for wiring logs queries correctly in the app, including safe SQL construction, query routing, and feature-flag-aware behavior. <!-- end of auto-generated comment: release notes by coderabbit.ai --> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This commit is contained in:
1 parent
719434a7fd
commit
3674f173a3
3 files changed
+412
No files matched your search
@@ -0,0 +1,223 @@
|
||||
---
|
||||
name: clickhouse-logs-queries
|
||||
description: >-
|
||||
Write, review, and migrate Supabase logs queries against the ClickHouse-backed
|
||||
`logs` table (the `logs.all.otel` analytics endpoint). Use this whenever a task
|
||||
involves Logs Explorer SQL, the `log_attributes` map, querying a log `source`
|
||||
(edge_logs, postgres_logs, auth_logs, etc.), translating an old BigQuery
|
||||
`cross join unnest(metadata)` logs query to ClickHouse, or wiring analytics log
|
||||
SQL in `apps/studio/data/logs` and `apps/studio/components/interfaces/Settings/Logs`.
|
||||
Reach for it even when the user just says "logs query", "Logs Explorer", or
|
||||
pastes a BigQuery logs query to convert, not only when they name ClickHouse.
|
||||
---
|
||||
|
||||
# Querying Supabase logs (ClickHouse)
|
||||
|
||||
Supabase logs live in a single ClickHouse `logs` table, served by the
|
||||
`logs.all.otel` analytics endpoint. Every log line from every part of the stack
|
||||
is one row in this table, tagged by a `source` column. This replaces the older
|
||||
BigQuery model, where each service had its own table and fields were reached
|
||||
through `cross join unnest(metadata)`.
|
||||
|
||||
Two kinds of work use this skill, and they share the same SQL model:
|
||||
|
||||
1. **Writing or reviewing a logs query** (in the Logs Explorer or anywhere a raw
|
||||
ClickHouse logs query is needed). Start here in this file.
|
||||
2. **Wiring a logs query in the Studio codebase** (branded analytics SQL, the
|
||||
endpoint picker, the OTEL query builders). Read
|
||||
[references/codebase-integration.md](references/codebase-integration.md).
|
||||
|
||||
If you are converting an existing BigQuery logs query, read
|
||||
[references/bigquery-migration.md](references/bigquery-migration.md) for the full
|
||||
translation table.
|
||||
|
||||
## The logs table
|
||||
|
||||
Each row has a small set of real columns. Everything specific to a service lives
|
||||
in `log_attributes`.
|
||||
|
||||
| Column | Type | Notes |
|
||||
| ---------------- | --------------------- | ----------------------------------------------------- |
|
||||
| `id` | `String` | Unique log identifier. |
|
||||
| `timestamp` | `DateTime64` (UTC) | When the log was produced. Order/compare it directly. |
|
||||
| `event_message` | `String` | The raw log line. |
|
||||
| `severity_text` | `String` | Log level, when the source sets one. |
|
||||
| `source` | `String` | The service the log came from. Always filter on this. |
|
||||
| `log_attributes` | `Map(String, String)` | Structured per-source fields, keyed by a dotted path. |
|
||||
|
||||
`timestamp` is formatted like `2026-06-22T09:34:06.215000` (ISO 8601, microsecond
|
||||
precision, no trailing `Z`). In the Logs Explorer the selected time range is
|
||||
applied for you, so you rarely need to write a `timestamp` filter by hand.
|
||||
|
||||
A minimal, well-formed query:
|
||||
|
||||
```sql
|
||||
select timestamp, event_message
|
||||
from logs
|
||||
where source = 'edge_logs'
|
||||
order by timestamp desc
|
||||
limit 100;
|
||||
```
|
||||
|
||||
## Sources
|
||||
|
||||
`source` selects the service. The common ones:
|
||||
|
||||
- `edge_logs` — API gateway requests and responses
|
||||
- `postgres_logs` — database statements and errors (also where pg_cron logs live)
|
||||
- `auth_logs` — authentication and authorization activity
|
||||
- `function_edge_logs` — edge function requests and responses
|
||||
- `function_logs` — `console` output from inside edge functions
|
||||
- `storage_logs` — object upload and retrieval activity
|
||||
- `realtime_logs` — Realtime client connections
|
||||
- `postgrest_logs`, `supavisor_logs`, `pgbouncer_logs` — mostly `id`, `timestamp`, `event_message`
|
||||
|
||||
The Logs Explorer **Field Reference** drawer lists every source and the fields it
|
||||
actually sets. When in doubt about a key, discover it from real data rather than
|
||||
guessing (see below).
|
||||
|
||||
## Reading fields from log_attributes
|
||||
|
||||
`log_attributes` maps a string key to a string value. Read a field with bracket
|
||||
access. There are no unnesting joins:
|
||||
|
||||
```sql
|
||||
select
|
||||
log_attributes['request.method'] as method,
|
||||
log_attributes['request.path'] as path,
|
||||
log_attributes['response.status_code'] as status
|
||||
from logs
|
||||
where source = 'edge_logs'
|
||||
```
|
||||
|
||||
The key keeps the dotted path that BigQuery expressed through nested structs, with
|
||||
the `metadata` root dropped: BigQuery `metadata.request.method` becomes
|
||||
`log_attributes['request.method']`. Keep the full prefix — `request.cf.country`
|
||||
is `log_attributes['request.cf.country']`, not `log_attributes['cf.country']`.
|
||||
|
||||
Common keys by source:
|
||||
|
||||
- `edge_logs`: `request.method`, `request.path`, `request.search`, `response.status_code`, `identifier`
|
||||
- `postgres_logs`: `parsed.error_severity`, `parsed.detail`, `parsed.hint`, `parsed.query`, `identifier`
|
||||
- `auth_logs`: `level`, `status`, `path`, `msg`, `error`
|
||||
- `function_edge_logs`: `response.status_code`, `request.method`, `request.pathname`, `function_id`, `execution_id`, `execution_time_ms`
|
||||
- `function_logs`: `event_type`, `function_id`, `execution_id`, `level`
|
||||
|
||||
### Numeric fields are strings
|
||||
|
||||
Map values are always strings. To compare or aggregate a numeric field, wrap it in
|
||||
`toInt32OrZero`, which returns `0` for missing or non-numeric values so it never
|
||||
errors on partial data:
|
||||
|
||||
```sql
|
||||
select count() as server_errors
|
||||
from logs
|
||||
where source = 'edge_logs'
|
||||
and toInt32OrZero(log_attributes['response.status_code']) between 500 and 599
|
||||
```
|
||||
|
||||
### Discover the keys a source sets
|
||||
|
||||
Read `mapKeys` from recent rows rather than guessing key names:
|
||||
|
||||
```sql
|
||||
select arrayJoin(mapKeys(log_attributes)) as key, count() as n
|
||||
from logs
|
||||
where source = 'postgres_logs'
|
||||
group by key
|
||||
order by n desc
|
||||
limit 100;
|
||||
```
|
||||
|
||||
`arrayJoin(mapKeys(...))` flattens the map keys into one row per key so you can
|
||||
rank them by frequency. (The Studio codebase does exactly this for the Field
|
||||
Reference drawer and to feed real keys to the AI rewrite.)
|
||||
|
||||
## ClickHouse vs BigQuery functions
|
||||
|
||||
These are the substitutions that trip people up most:
|
||||
|
||||
| Need | BigQuery | ClickHouse |
|
||||
| ------------------ | ----------------------------- | -------------------------------------------- |
|
||||
| Count rows | `count(*)` | `count()` |
|
||||
| Regex match | `regexp_contains(x, 'p')` | `match(x, 'p')` |
|
||||
| Substring match | `x like '%p%'` | `x ilike '%p%'` (case-insensitive) or `like` |
|
||||
| Numeric coercion | `cast(x as int64)` | `toInt32OrZero(x)` |
|
||||
| Read the timestamp | `cast(timestamp as datetime)` | `timestamp` (use the column directly) |
|
||||
| Map keys | n/a (used unnest) | `mapKeys(log_attributes)` |
|
||||
|
||||
The `logs.all.otel` analytics endpoint (and the Logs Explorer on top of it)
|
||||
rejects `count(*)` and `select *` — use `count()` and list the columns you need.
|
||||
(Raw ClickHouse supports both; this is a constraint of the logs query surface.)
|
||||
|
||||
## Best practices
|
||||
|
||||
These keep queries correct and cheap. Log tables are large; an unbounded scan
|
||||
reads far more data than you need.
|
||||
|
||||
- **Always include a `LIMIT`.** Even for aggregates while you iterate.
|
||||
- **Always filter by `source`.** It scopes the query to one service.
|
||||
- **Keep the time range tight.** A smaller window returns results faster.
|
||||
- **Filter on the real columns** (`source`, `timestamp`) before reaching into
|
||||
`log_attributes`.
|
||||
- **Order by `timestamp desc`** to see the most recent logs first.
|
||||
- **Use `count()`**, not `count(*)` or `select *`.
|
||||
|
||||
## Worked examples
|
||||
|
||||
Requests by status code:
|
||||
|
||||
```sql
|
||||
select
|
||||
toInt32OrZero(log_attributes['response.status_code']) as status,
|
||||
count() as count
|
||||
from logs
|
||||
where source = 'edge_logs'
|
||||
group by status
|
||||
order by count desc
|
||||
limit 50
|
||||
```
|
||||
|
||||
Auth errors:
|
||||
|
||||
```sql
|
||||
select timestamp, event_message, log_attributes['msg'] as message
|
||||
from logs
|
||||
where source = 'auth_logs'
|
||||
and log_attributes['level'] in ('error', 'fatal')
|
||||
order by timestamp desc
|
||||
limit 100
|
||||
```
|
||||
|
||||
Search the raw message:
|
||||
|
||||
```sql
|
||||
select timestamp, event_message
|
||||
from logs
|
||||
where source = 'postgres_logs'
|
||||
and event_message ilike '%deadlock%'
|
||||
order by timestamp desc
|
||||
limit 100
|
||||
```
|
||||
|
||||
Postgres errors grouped by severity (the canonical unnest-to-map conversion):
|
||||
|
||||
```sql
|
||||
select log_attributes['parsed.error_severity'] as severity, count() as count
|
||||
from logs
|
||||
where source = 'postgres_logs'
|
||||
and log_attributes['parsed.error_severity'] in ('ERROR', 'FATAL', 'PANIC')
|
||||
group by severity
|
||||
order by count desc
|
||||
limit 100
|
||||
```
|
||||
|
||||
## When the user pastes a BigQuery query
|
||||
|
||||
Convert it rather than running it as-is. The mechanical steps (drop the
|
||||
per-service table for `from logs where source = ...`, remove every
|
||||
`cross join unnest(...)`, rewrite unnest-alias columns as `log_attributes['...']`
|
||||
lookups, swap the functions above) are spelled out with a full before/after in
|
||||
[references/bigquery-migration.md](references/bigquery-migration.md). The Logs
|
||||
Explorer also has a built-in **Rewrite to ClickHouse** action that does this with
|
||||
AI; point users to it for one-off conversions in the dashboard.
|
||||
@@ -0,0 +1,93 @@
|
||||
# Migrating a BigQuery logs query to ClickHouse
|
||||
|
||||
The old logs engine gave each service its own table and exposed structured fields
|
||||
through repeated `cross join unnest(...)` over a nested `metadata` column. The
|
||||
ClickHouse engine has one `logs` table and a flat `log_attributes` map. Conversion
|
||||
is mechanical once you internalize the mapping.
|
||||
|
||||
## The five steps
|
||||
|
||||
1. **Replace the table with `logs` plus a `source` filter.** The old table name is
|
||||
the `source` value: `from postgres_logs as t` becomes
|
||||
`from logs where source = 'postgres_logs'`. Never select from a per-service
|
||||
table name (`postgres_logs`, `edge_logs`, ...) on the ClickHouse engine.
|
||||
- Exception: pg_cron logs live under `source = 'postgres_logs'`.
|
||||
|
||||
2. **Remove every unnest join.** Delete `cross join unnest(metadata) as m`,
|
||||
`cross join unnest(m.parsed) as p`, `left join unnest(...) on true`, and so on.
|
||||
They have no equivalent — the data is already flat in the map.
|
||||
|
||||
3. **Rewrite every unnest-alias column as a map lookup.** A field taken off
|
||||
`unnest(metadata)` becomes `log_attributes['field']`. A field off a nested
|
||||
struct like `unnest(m.parsed)` becomes `log_attributes['parsed.field']`: keep
|
||||
the struct name as a dotted prefix, drop the `metadata` root and every alias.
|
||||
|
||||
4. **Wrap numeric fields in `toInt32OrZero(...)`** before comparing or aggregating
|
||||
them — map values are strings.
|
||||
|
||||
5. **Swap BigQuery functions for ClickHouse ones** (see the table below).
|
||||
|
||||
Preserve the original select list, filters, group by, order by, and limit intent
|
||||
throughout.
|
||||
|
||||
## Function substitutions
|
||||
|
||||
| Need | BigQuery | ClickHouse |
|
||||
| ---------------- | ----------------------------- | --------------------------- |
|
||||
| Count rows | `count(*)` | `count()` |
|
||||
| Regex match | `regexp_contains(x, 'p')` | `match(x, 'p')` |
|
||||
| Substring match | `x like '%p%'` | `x ilike '%p%'` or `like` |
|
||||
| Numeric coercion | `cast(x as int64)` | `toInt32OrZero(x)` |
|
||||
| Timestamp value | `cast(timestamp as datetime)` | `timestamp` |
|
||||
| All columns | `select *` | list the columns explicitly |
|
||||
|
||||
## Full before/after
|
||||
|
||||
BigQuery:
|
||||
|
||||
```sql
|
||||
select count(t.timestamp) as count, p.error_severity
|
||||
from
|
||||
postgres_logs as t
|
||||
cross join unnest(metadata) as m
|
||||
cross join unnest(m.parsed) as p
|
||||
where p.error_severity in ('ERROR', 'FATAL', 'PANIC')
|
||||
group by p.error_severity
|
||||
order by count desc
|
||||
limit 100;
|
||||
```
|
||||
|
||||
ClickHouse:
|
||||
|
||||
```sql
|
||||
select count() as count, log_attributes['parsed.error_severity'] as error_severity
|
||||
from logs
|
||||
where source = 'postgres_logs'
|
||||
and log_attributes['parsed.error_severity'] in ('ERROR', 'FATAL', 'PANIC')
|
||||
group by log_attributes['parsed.error_severity']
|
||||
order by count desc
|
||||
limit 100
|
||||
```
|
||||
|
||||
Notice: `count(t.timestamp)` became `count()`, the two unnest joins are gone,
|
||||
`p.error_severity` became `log_attributes['parsed.error_severity']`, and the
|
||||
`from`/`where` targets the single table.
|
||||
|
||||
## Getting the keys right
|
||||
|
||||
The most common conversion mistake is dropping a dotted prefix — writing
|
||||
`log_attributes['x_real_ip']` when the real key is
|
||||
`log_attributes['request.headers.x_real_ip']`, or `log_attributes['cf.country']`
|
||||
instead of `log_attributes['request.cf.country']`. Do not invent or shorten keys.
|
||||
When unsure, discover the real keys from data:
|
||||
|
||||
```sql
|
||||
select arrayJoin(mapKeys(log_attributes)) as key, count() as n
|
||||
from logs
|
||||
where source = 'edge_logs'
|
||||
group by key
|
||||
order by n desc
|
||||
limit 200;
|
||||
```
|
||||
|
||||
Then map each old nested path to the matching key exactly as it appears.
|
||||
@@ -0,0 +1,96 @@
|
||||
# Wiring a logs query in the Studio codebase
|
||||
|
||||
This covers writing analytics log SQL inside `apps/studio` so it is safe, routed
|
||||
to the right endpoint, and consistent with the existing OTEL builders. Read this
|
||||
when you are editing TypeScript that builds or runs a logs query, not when you are
|
||||
just writing a query in the Logs Explorer UI.
|
||||
|
||||
## Branded SQL: never concatenate user input
|
||||
|
||||
All analytics log SQL must be a `SafeLogSqlFragment`, built with the helpers in
|
||||
`apps/studio/data/logs/safe-analytics-sql.ts`. This is enforced by eslint, and the
|
||||
branding is what keeps interpolated values from becoming injection. The key
|
||||
exports:
|
||||
|
||||
- `safeSql\`...\``— a tagged template that only accepts`SafeLogSqlFragment`interpolations. Plain strings (and Postgres-branded`SafeSqlFragment`) are
|
||||
rejected at compile time, so you cannot accidentally drop a raw value in.
|
||||
- `analyticsLiteral(value)` — turns a string/number/boolean into a safely escaped
|
||||
literal fragment (single quotes and backslashes are escaped). Use it for every
|
||||
dynamic value, especially the `source`.
|
||||
- `joinSqlFragments(fragments, separator)` — joins already-safe fragments with a
|
||||
fixed structural separator (`' and '`, `', '`, etc.).
|
||||
- `keyword(value, allowed)` — resolves a value against a compile-time allow-list of
|
||||
fragments (e.g. an `AND`/`OR` operator). Returns the allow-listed fragment, never
|
||||
the raw input.
|
||||
- `quotedIdent(value)` — backtick-quotes a dotted identifier path after validating
|
||||
each segment.
|
||||
|
||||
There is intentionally no exported "raw" escape hatch. Compose with `safeSql` plus
|
||||
these helpers.
|
||||
|
||||
```ts
|
||||
import { analyticsLiteral, safeSql } from 'data/logs/safe-analytics-sql'
|
||||
|
||||
const source = 'edge_logs'
|
||||
const sql = safeSql`
|
||||
select timestamp, event_message
|
||||
from logs
|
||||
where source = ${analyticsLiteral(source)}
|
||||
order by timestamp desc
|
||||
limit 100
|
||||
`
|
||||
```
|
||||
|
||||
## Pick the endpoint and builder by flag
|
||||
|
||||
The ClickHouse path is gated by the `otelLegacyLogs` PostHog flag
|
||||
(`useFlag('otelLegacyLogs')` from `common`). Keep the BigQuery path working when
|
||||
the flag is off. Two helpers in `apps/studio/data/logs/logs-endpoint.ts` express
|
||||
the split:
|
||||
|
||||
- `logsAllEndpointUrl(useOtel)` — returns the `logs.all.otel` endpoint when
|
||||
`useOtel`, otherwise the legacy `logs.all` endpoint.
|
||||
- `pickLogsQueryBuilder(useOtel, otel, bq)` — picks between an OTEL builder and a
|
||||
BigQuery builder while preserving the type.
|
||||
|
||||
```ts
|
||||
const useOtel = useFlag('otelLegacyLogs')
|
||||
const builder = pickLogsQueryBuilder(useOtel, genDefaultQueryOtel, genDefaultQuery)
|
||||
const endpoint = logsAllEndpointUrl(useOtel)
|
||||
// include { otel: useOtel } in the React Query key so the two paths cache separately
|
||||
```
|
||||
|
||||
Run the fragment through `executeAnalyticsSql` (`apps/studio/data/logs/execute-analytics-sql.ts`)
|
||||
against that endpoint.
|
||||
|
||||
## Follow the existing OTEL builders
|
||||
|
||||
When you need a new query shape, mirror the generators in
|
||||
`apps/studio/components/interfaces/Settings/Logs/Logs.utils.otel.ts` rather than
|
||||
inventing a parallel style. They already encode the conventions:
|
||||
|
||||
- `genDefaultQueryOtel`, `genCountQueryOtel`, `genChartQueryOtel`,
|
||||
`genSingleLogQueryOtel` — the row/count/chart/single-log builders. They select
|
||||
the real columns plus source-specific `log_attributes[...]` lookups aliased to
|
||||
the leaf names the renderers expect.
|
||||
- `mapOtelPreviewRow`, `mapOtelSingleLogToLegacy`, `otelTimestampToMicros` — the
|
||||
JS normalization layer. `timestamp` must end up a microsecond number for the
|
||||
pagination cursor and the renderers, so reuse `otelTimestampToMicros` rather
|
||||
than parsing the ISO string yourself.
|
||||
|
||||
The map-key discovery hook (`apps/studio/data/logs/otel-log-keys-query.ts`,
|
||||
`fetchOtelLogKeys` / `useOtelLogKeysQuery`) is the canonical example of a small,
|
||||
correctly-branded OTEL query — read it before writing a new one.
|
||||
|
||||
## Checklist
|
||||
|
||||
- [ ] Every dynamic value goes through `analyticsLiteral` (or another sanitizer),
|
||||
never string concatenation.
|
||||
- [ ] The query filters by `source` and includes a `LIMIT`.
|
||||
- [ ] Numeric `log_attributes` values are wrapped in `toInt32OrZero`.
|
||||
- [ ] The endpoint and builder are chosen with `logsAllEndpointUrl` /
|
||||
`pickLogsQueryBuilder` off `useFlag('otelLegacyLogs')`.
|
||||
- [ ] The React Query key distinguishes the OTEL and BigQuery paths.
|
||||
- [ ] `timestamp` is normalized to micros for any row consumed by the table/cursor.
|
||||
- [ ] There is a unit test asserting the generated SQL string (see
|
||||
`Logs.utils.otel.test.ts` and `safe-analytics-sql.test.ts` for the pattern).
|
||||
Reference in new issue
Block a user