diff --git a/apps/docs/docs/ref/self-hosting-analytics/introduction.mdx b/apps/docs/docs/ref/self-hosting-analytics/introduction.mdx index 2c2c31ab4be..bc41e3509a2 100644 --- a/apps/docs/docs/ref/self-hosting-analytics/introduction.mdx +++ b/apps/docs/docs/ref/self-hosting-analytics/introduction.mdx @@ -15,8 +15,14 @@ hideTitle: true The Supabase Analytics server is a Logflare self-hostable instance that manages the ingestion and query pipelines for searching and aggregating structured analytics events. -When self-hosting the Analytics server, the full logging experience matching that of the Supabase Platform is available in the Studio instance, allowing for an integrated and enhanced development experience. -However, it's important to note that certain [differences](#differences) may arise due to the platform's infrastructure. +When running the Analytics server, the logging experience matching that of the Supabase Platform is available in Studio. However, log content and metadata can differ from the Supabase Platform - see **[Differences from the Supabase platform](#differences-from-the-supabase-platform)** below. + +Locally running logs and analytics are supported in two separate setups: + +- **[Local development & CLI](/docs/guides/local-development)** - a local Supabase stack managed with the Supabase CLI (`supabase start`). +- **[Self-hosted Supabase](/docs/guides/self-hosting)** - the full stack on your own infrastructure with Docker Compose. Logs & Analytics are **optional** and are not started by default. + +The sections below cover backend options and Logflare configuration. For self-hosted Supabase, see the [installation and configuration](/docs/guides/self-hosting/docker) guide. @@ -24,158 +30,237 @@ All Logflare technical documentation is available at https://docs.logflare.app. -## Backends Supported +## Supported backends -The Analytics server supports either **Postgres** or **BigQuery** as the backend. The `supabase-cli` experience uses the Postgres backend out-of-the-box. However, the Supabase Platform uses the BigQuery backend for storing all platform logs. +The Analytics server supports either **Postgres** or **BigQuery** as the backend. Both [Local development & CLI](/docs/guides/local-development) and [self-hosted Supabase](/docs/guides/self-hosting) use the Postgres backend out of the box. When using the BigQuery backend, a BigQuery dataset is created in the provided Google Cloud project, and tables are created for each source. Log events are streamed into each table, and all queries generated by Studio or by the Logs Explorer are executed against the BigQuery API. This backend requires internet access to work, and cannot be run fully locally. -When using the Postgres backend, tables are created for each source within the provided schema (for supabase-cli, this would be `_analytics`). Log events received by Logflare are inserted directly into the respective tables. All BigQuery-dialect SQL queries from Studio will be handled by a translation layer within the Analytics server. This translation layer translates the query to Postgres dialect, and then executes it against the Postgres database. +When using the Postgres backend, tables are created for each source within the provided schema. Log events received by Logflare are inserted directly into the respective tables. All BigQuery-dialect SQL queries from Studio are handled by a translation layer **within the Analytics server**. This translation layer translates the query to Postgres dialect, and then executes it against the Postgres database. The Postgres backend is not yet optimized for a high volume of inserts, or for heavy query usage. Today the translation layer only handles a limited subset of the BigQuery dialect. As such, the [Log Explorer](https://supabase.com/docs/guides/platform/logs#logs-explorer) may produce errors for more advanced queries when using the Postgres Backend. -## Getting Started +## Getting started -The Postgres backend is recommended when familiarizing and experimenting with self-hosting Supabase. For production, we recommend using the BigQuery backend. See [production recommendations](#production-recommendations) for more information. +The Postgres backend is recommended when familiarizing and experimenting. For production self-hosting, we recommend the BigQuery backend. See the [notes](#recommendations-for-production) below for more information. -To set up logging in self-hosted Supabase, see the [docker-compose example](https://github.com/supabase/supabase/tree/master/docker). -Two compose services are required: Logflare, and Vector. Logflare is the HTTP Analytics server, while Vector is the logging pipeline to route all compose services' syslog to the Logflare sever. +### Local development & CLI -Regardless of the backend chosen, the following environment variables **must** be set for the `supabase/logflare` docker image: +With the Supabase CLI, run `supabase start` to bring up a local stack. Logs use the Postgres backend by default and are stored in the `_analytics` schema. Docker socket access may need manual configuration on some platforms - see the [Access your project's services](/docs/guides/local-development/cli/getting-started#access-your-projects-services) (the Analytics tab). -- `LOGFLARE_SINGLE_TENANT=true`: The feature flag for enabling single tenant mode for Logflare. Must be set to `true` -- `LOGFLARE_SUPABASE_MODE=true`: The feature flag for seeding Supabase-related data. Must be set to `true` + -For all other configuration environment variables, please refer to the [Logflare self-hosting documentation](https://docs.logflare.app/self-hosting/#configuration). - -## Postgres Backend Setup - -The [example docker-compose](https://github.com/supabase/supabase/tree/master/docker) uses the Postgres backend out of the box. - -```bash -# clone the supabase/supabase repo, and run the following -cd docker -docker compose -f docker-compose.yml up -``` - -### Configuration and Requirements - -- `supabase/logflare:1.4.0` or above -- Relevant environment variables: - - `POSTGRES_BACKEND_URL` : Required. The connection string to the Postgres database. - - `POSTGRES_BACKEND_SCHEMA` : Optional. Allows customization of the schema used to scope all backend operations within the database. - -## BigQuery Backend Setup - -The BigQuery backend is a more robust and scalable backend option that is battle-tested and production ready. Use this backend if you intend to have heavy logging usage and require advanced querying features such as the Logs Explorer. - -### Configuration and Requirements - -The requirements are as follows after creating the project: - -- Google Cloud project with billing enabled -- Project ID -- Project number -- A service account key. - - - -You must enable billing on your Google Cloud project, as a valid billing account is required for streaming inserts. +The Supabase CLI is intended for local development only. The local analytics service uses predefined, publicly known access tokens and has no protection against external traffic. Never run the CLI stack on a publicly accessible or untrusted network. -#### Setting up BigQuery Service Account +### Self-hosted Supabase -The service account used must have sufficient permissions to insert into your Google Cloud BigQuery. Ensure that the service account has either: +Self-hosted Supabase does **not** include Logs & Analytics in the default Docker Compose configuration. You must first install and secure the base stack - generate secrets, configure URLs, and set a Studio password - as described in [Self-Hosting with Docker](/docs/guides/self-hosting/docker). -- BigQuery Admin role; or -- The following permissions: - - `bigquery.datasets.create` - - `bigquery.datasets.get` - - `bigquery.datasets.getIamPolicy` - - `bigquery.datasets.update` - - `bigquery.jobs.create` - - `bigquery.routines.create` - - `bigquery.routines.update` - - `bigquery.tables.create` - - `bigquery.tables.delete` - - `bigquery.tables.get` - - `bigquery.tables.getData` - - `bigquery.tables.update` - - `bigquery.tables.updateData` + -You can create the service account via the web console or `gcloud` CLI, as per the [Google Cloud documentation](https://cloud.google.com/iam/docs/keys-create-delete). In the web console, you can create the key by navigating to IAM > Service Accounts > Actions (dropdown) > Manage Keys +**Do not** start the stack with placeholder values from `.env.example`. -We recommend setting the BigQuery Admin role, as it simplifies permissions setup. + -#### Download the Service Account Keys +After the base stack is configured, enable Logs & Analytics from your project directory (the directory containing `docker-compose.yml`): -After the service account is created, you will need to create a key for the service account. This key will sign the JWTs for API requests that the Analytics server makes with BigQuery. This can be done through the IAM section in Google Cloud console. - -#### Docker Image Configuration - -Using the example [self-hosting stack based on docker-compose](https://github.com/supabase/supabase/tree/master/docker), you include the logging related services using the following command - -1. Update the `.env.example` file with the necessary environment variables. - -- `GOOGLE_PROJECT_ID` -- `GOOGLE_PROJECT_NUMBER` - -2. Place your Service Account key in your present working directory with the filename `gcloud.json`. -3. On `docker-compose.yml`, uncomment the block section below the commentary `# Uncomment to use Big Query backend for analytics` -4. On `docker-compose.yml`, comment the block section below the commentary `# Comment variables to use Big Query backend for analytics` - -Thereafter, you can start the example stack using the following command: - -```bash -# assuming you clone the supabase/supabase repo. -cd docker -docker compose -f docker-compose.yml +```sh +sh run.sh config add logs && \ +sh run.sh start ``` -#### BigQuery Datset Storage Location +This layers [`docker-compose.logs.yml`](https://github.com/supabase/supabase/blob/master/docker/docker-compose.logs.yml) on top of the base configuration and starts two additional services (`analytics` and `vector`). -Currently, all BigQuery datasets stored and managed by Analytics, whether via CLI or self-hosted, will default to the US region. +See [Enabling analytics](/docs/guides/self-hosting/docker#enabling-analytics) in the self-hosting guide for more context. -## Vector Usage +### Required Logflare configuration -In the Docker Compose example, Vector is used for the logging pipieline, where log events are forwarded to the Analytics API for ingestion. +The following environment variables are always required for the `supabase/logflare` image: + +```yaml name=docker-compose.logs.yml +analytics: + environment: + # Enable single-tenant mode for Logflare + LOGFLARE_SINGLE_TENANT: 'true' + # Seed Supabase-related metadata + LOGFLARE_SUPABASE_MODE: 'true' +``` + +For all other configuration environment variables, refer to the [Logflare self-hosting documentation](https://docs.logflare.app/self-hosting/#configuration). + +## Using the Postgres backend + +### Local development & CLI + +The CLI uses the Postgres backend by default. No additional configuration is needed. + +### Self-hosted Supabase + +The [`docker-compose.logs.yml`](https://github.com/supabase/supabase/blob/master/docker/docker-compose.logs.yml) override uses the Postgres backend by default. + +#### Configuration and requirements + +The following environment variables control the Postgres backend (set in `docker-compose.logs.yml`): + +- `POSTGRES_BACKEND_URL`: Required. Connection string to the Postgres database. +- `POSTGRES_BACKEND_SCHEMA`: Optional. Schema used to store log data (default: `_analytics`). + +## Using the BigQuery backend + +The BigQuery backend is a more robust and scalable backend option that is battle-tested and production ready. Use this backend if you intend to have heavy logging usage and require advanced querying features such as the Logs Explorer. + +### Requirements + +Using the BigQuery backend requires: + +- A Google Cloud project with billing enabled +- The project ID +- The project number +- A service account JSON key + +A valid billing account is required for BigQuery streaming inserts. + +### Google Cloud Console setup + +1. Open [Google Cloud Console](https://console.cloud.google.com/) and select or create the project you want to use for analytics. +2. Confirm **billing is enabled** on that project ([Billing](https://console.cloud.google.com/billing) - link the project to a billing account if needed). +3. Go to **IAM & Admin > Service Accounts** and click **Create service account** (for example, `supabase-analytics`). +4. On **Permissions**, grant **BigQuery Admin** to keep setup simple. Skip **Principals with access**. Alternatively, grant only the specific permissions listed below instead of the admin role. +5. Open the service account, then **Keys** (or **Actions > Manage keys** from the list view). Click **Add key > Create new key**, choose **JSON**, and download the key file. +6. Go to **APIs & Services > Library**, search for **Cloud Resource Manager API**, open it, and click **Enable**. Confirm the correct project is selected in the project picker at the top. You can also enable it directly: `https://console.developers.google.com/apis/api/cloudresourcemanager.googleapis.com/overview?project=YOUR_PROJECT_ID` (replace `YOUR_PROJECT_ID`). +7. Go to **Cloud overview > Dashboard** (home dashboard for the project) and copy the **Project ID** and **Project number**. Confirm the project picker at the top shows the same project you configured above. +8. Save the downloaded JSON key as `gcloud.json` in your self-hosted Supabase project directory (the directory that contains `docker-compose.yml`. + + + +Do not commit `gcloud.json` to version control. Add it to `.gitignore` if needed. + + + + + +If you prefer not to use the **BigQuery Admin** role, the service account needs at least the following permissions: + +- `bigquery.datasets.create`, `bigquery.datasets.get`, `bigquery.datasets.getIamPolicy`, `bigquery.datasets.update` +- `bigquery.jobs.create` +- `bigquery.routines.create`, `bigquery.routines.update` +- `bigquery.tables.create`, `bigquery.tables.delete`, `bigquery.tables.get`, `bigquery.tables.getData`, `bigquery.tables.update`, `bigquery.tables.updateData` + + + +### Self-hosted Supabase configuration + +Complete [Self-Hosting with Docker](/docs/guides/self-hosting/docker) first. Enable the logs override, then switch the `analytics` service in [`docker-compose.logs.yml`](https://github.com/supabase/supabase/blob/master/docker/docker-compose.logs.yml) from Postgres to BigQuery. + +**1. Set Google Cloud variables in `.env`** + +Add your project values under the Analytics section (do not leave the `.env.example` placeholders): + +```sh +############ +# Logs and Analytics +############ + +# Google Cloud Project details +GOOGLE_PROJECT_ID=your-project-id +GOOGLE_PROJECT_NUMBER=123456789012 +``` + +**2. Mount the service account key and switch the analytics backend** + +In `docker-compose.logs.yml`, under the `analytics` service: + +- Remove or comment out `POSTGRES_BACKEND_URL` and `POSTGRES_BACKEND_SCHEMA`. +- Uncomment `GOOGLE_PROJECT_ID` and `GOOGLE_PROJECT_NUMBER`. +- Uncomment the `volumes` block and bind-mount `gcloud.json` from your project directory. + +Example: + +```yml name=docker-compose.logs.yml +analytics: + environment: + GOOGLE_PROJECT_ID: ${GOOGLE_PROJECT_ID} + GOOGLE_PROJECT_NUMBER: ${GOOGLE_PROJECT_NUMBER} + volumes: + - ./gcloud.json:/opt/app/rel/logflare/bin/gcloud.json:ro,z +``` + +**3. Start the stack** + +From your project directory: + +```sh +sh run.sh config add logs && \ +sh run.sh start +``` + +## Log collection and routing with Vector + +In the self-hosted Docker Compose setup, [Vector](https://github.com/vectordotdev/vector) is used for the logging pipeline when the `docker-compose.logs.yml` override is enabled. Log events are forwarded to the Analytics API for ingestion. Please refer to the [Vector configuration file](https://github.com/supabase/supabase/blob/master/docker/volumes/logs/vector.yml) when customizing your own setup. -You **must** ensure that the payloads matches the expected event schema structure. Without the correct structure, it would cause the Studio Logs UI features to break. +You must ensure that each payload matches the event shape produced by the Vector transforms. Logs Explorer in Studio expects the fields below when querying Analytics sources: -## Differences from Platform +| Field | Description | +| --------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | +| `event_message` | The log message displayed in Studio. | +| `timestamp` | The event timestamp. If omitted, Logflare assigns an ingestion timestamp. | +| `metadata` | A JSON object with service-specific fields used by the Logs UI. For example, API logs include request and response metadata, while database logs include parsed Postgres severity metadata. | -API logs rely on Kong instead of the Supabase Cloud API Gateway. Logs from Kong are not enriched with platform-only data. +When adding a custom Vector source, map your service logs to this shape before sending them to `http://analytics:4000/api/logs?source_name=`. Reuse the existing source names in `vector.yml` if you want the logs to appear in the built-in Studio Logs panels. -Within the self-hosted setup, all logs are routed to Logflare via Vector. As Kong routes API requests to PostgREST, self-hosted or local deployments will result in Kong request logs instead. -This would result in differences in the log event metadata between self-hosted API requests and Supabase Platform requests. +## Differences from the Supabase platform -## Production Recommendations +API logs rely on the locally running API gateway (Kong) instead of the Supabase Cloud API Gateway. Logs from Kong are not enriched with platform-only data. -To self-host in a production setting, we recommend performing the following for a better experience. +Within self-hosted Supabase (with the logs override enabled), container logs are routed to Logflare via Vector. As Kong routes API requests to PostgREST, self-hosted and [local development & CLI](/docs/guides/local-development) deployments produce Kong request logs rather than the platform API gateway logs. -### Ensure that Logflare is behind a firewall and restrict all network access to it besides safe requests. +This results in differences in log event metadata between your own infrastructure and the Supabase Platform. -Self-hosted Logflare has UI authentication disabled and is intended for exposure to the internet. We recommend restricting access to the dashboard, accessible at the `/dashboard` path. -If dashboard access is required for managing sources, we recommend having an authentication layer, such as a VPN. +## Recommendations for production -### Use a different Postgres Database to store Logflare data. +When running a self-hosted instance of Logflare in production, we recommend the following. -Logflare requires a Postgres database to function. However, if there is an issue with you self-hosted Postgres service, you would not be able to debug it as it would also bring Logflare down together. +### Set unique access tokens -The self-hosted example is only used as a minimal example on running the entire stack, however it is not recommended to use the same database server for both production and observability. + -### Use Big Query as the Logflare Backend +`LOGFLARE_PUBLIC_ACCESS_TOKEN` and `LOGFLARE_PRIVATE_ACCESS_TOKEN` **must** be set to unique, randomly generated values. Never run with the placeholder or default values. -The current Postgres Ingestion backend isn't optimized for production usage. We recommend using Big Query for more heavy use cases. + + +In a self-hosted Supabase configuration, these tokens are generated automatically when you run `sh utils/generate-keys.sh` during initial setup. + +### Restrict network access to Logflare + +Self-hosted Logflare has UI authentication disabled. + + + +The `/dashboard` path is publicly accessible by default and **must** be restricted at the network level. If you need dashboard access for managing sources, place it behind a VPN or other secure access layer. + + + +Direct access to the Logflare's container port 4000 is intentionally not exposed in the default self-hosted Supabase configuration - Logflare is only reachable through the API gateway (Kong). Do not re-expose this port directly unless required. + +### Use a different Postgres Database to store Logflare data + +Logflare requires a Postgres database to function. However, if there is an issue with you self-hosted Postgres service, you would not be able to debug it as it would also bring down Logflare. + +The self-hosted Docker example is a minimal reference for running the entire stack; it is not recommended to use the same Postgres instance for both production workloads and observability data. + +### Use BigQuery as the Logflare backend + +The current Postgres Ingestion backend isn't optimized for production usage. We recommend using BigQuery for more heavy use cases. **We recommend using the BigQuery backend for production environments as it offers better scaling and querying/debugging experiences.** -### Rotate Encryption Keys Regularly +### Rotate encryption keys regularly -The Logflare server uses the a Base64 encryption key set on the `LOGFLARE_DB_ENCRYPTION_KEY` environment variable to perform encryption at rest for sensitive database columns. +The Logflare server uses the a Base64 encryption key optionally set via the `LOGFLARE_DB_ENCRYPTION_KEY` environment variable to perform encryption at rest for sensitive database columns. To perform encryption key rotation, move the retired key to the `LOGFLARE_DB_ENCRYPTION_KEY_RETIRED` environment variable, and replace the `LOGFLARE_DB_ENCRYPTION_KEY` environement variable with the new key. Perform a server restart and check `info` logs for the migration to be detected and performed.