diff --git a/README.md b/README.md index 94a4ffa..c60171e 100644 --- a/README.md +++ b/README.md @@ -31,6 +31,7 @@ Head over to the [ATProto Touchers Discord](https://discord.atprotocol.dev/) to - [Setting up SMTP](#setting-up-smtp) - [Common SMTP issues](#common-smtp-issues) - [Logging](#logging) + - [Monitoring and metrics](#monitoring-and-metrics) - [Updating your PDS](#updating-your-pds) - [Environment Variables](#environment-variables) - [Migrating your PDS](#migrating-your-pds) @@ -319,6 +320,22 @@ You can also change the minimum level of logs to be printed (default: `info`): LOG_LEVEL=debug ``` +### Monitoring and metrics + +The PDS can report metrics over [OpenTelemetry](https://opentelemetry.io/): host-level things like CPU, memory and disk, alongside PDS-level things like accounts created, sign-ins, OAuth grants, and XRPC request rate and latency. + +The [`monitoring/`](./monitoring) directory in this repo contains a self-contained Prometheus + Grafana + node_exporter stack and a ready-made Grafana dashboard. It is entirely optional, runs separately from the main PDS stack, and is not affected by `pdsadmin update`. + +```bash +curl -sL https://github.com/bluesky-social/pds/archive/refs/heads/main.tar.gz \ + | tar xz --strip-components=1 pds-main/monitoring +cd monitoring && docker compose up --detach +``` + +You then add some `OTEL_*` variables to `/pds/pds.env` and restart the PDS. Everything binds to `127.0.0.1`, so you can reach Grafana over an SSH tunnel rather than opening ports. See [monitoring/README.md](./monitoring/README.md) for the full walkthrough of both paths. + +If you already run Prometheus and Grafana, you don't need the compose file — point the PDS at your own OTLP endpoint and import the dashboard JSON. + ### Updating your PDS It is recommended that you keep your PDS up to date with new versions. This repo sets up [watchtower](https://github.com/nicholas-fedor/watchtower), which handles automatic updates. You can also use the `pdsadmin` tool to manually update your PDS. diff --git a/monitoring/README.md b/monitoring/README.md new file mode 100644 index 0000000..44924d8 --- /dev/null +++ b/monitoring/README.md @@ -0,0 +1,132 @@ +# Monitoring a PDS + +A ready-made Grafana dashboard for a self-hosted PDS, covering both host health +(CPU, memory, disk, network) and PDS activity (accounts, sessions, OAuth grants, +XRPC request rate and latency). + +This is entirely optional and completely separate from the main PDS stack. It +does not modify `/pds/compose.yaml` and is not touched by `pdsadmin update`. + +## How metrics get out of the PDS + +The PDS speaks [OpenTelemetry](https://opentelemetry.io/). It **pushes** metrics +over OTLP rather than exposing a `/metrics` endpoint to be scraped. + +Prometheus v3 can receive OTLP directly, so this stack is just three containers +with no OpenTelemetry Collector in between. + +## Quick start + +Grab this directory and start the stack: + +```bash +curl -sL https://github.com/bluesky-social/pds/archive/refs/heads/main.tar.gz \ + | tar xz --strip-components=1 pds-main/monitoring +cd monitoring && docker compose up --detach +``` + +Then tell the PDS where to send metrics. Add to `/pds/pds.env`: + +```bash +OTEL_SERVICE_NAME=pds +OTEL_EXPORTER_OTLP_METRICS_ENDPOINT=http://localhost:9090/api/v1/otlp/v1/metrics +OTEL_EXPORTER_OTLP_METRICS_PROTOCOL=http/protobuf +OTEL_METRIC_EXPORT_INTERVAL=15000 +OTEL_SEMCONV_STABILITY_OPT_IN=http +``` + +and restart: + +```bash +sudo systemctl restart pds +``` + +Two notes on those variables: + +- Setting only `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT` rather than the generic + `OTEL_EXPORTER_OTLP_ENDPOINT` is intentional. The PDS enables exactly the + OTel signals you configure, so this turns on metrics and leaves traces and + logs off. Setting the generic endpoint would point traces and logs at + Prometheus too, which cannot accept them. +- `OTEL_SEMCONV_STABILITY_OPT_IN=http` opts into the stable HTTP semantic + conventions (`http.server.request.duration`, in seconds). The dashboard is + built against those names. Without it, older HTTP metric names are emitted + and the request-rate and latency panels stay empty. + +## Reaching Grafana + +Everything binds to `127.0.0.1` and none of it is exposed to the internet. Do +not open these ports on your cloud firewall. Use an SSH tunnel: + +```bash +ssh -L 3001:localhost:3001 you@your-pds-host +``` + +Then open and log in with `admin` / `admin`. Grafana +will ask you to change the password on first login. + +Grafana runs on **3001** because the PDS itself owns port 3000. + +The **PDS Overview** dashboard is provisioned automatically, in a folder named +PDS. It is read-only; to customize it, use "Save as" to make your own copy, or +edit `dashboards/pds-overview.json` and restart Grafana. + +## Already running Prometheus and Grafana? + +In this case, don't use `compose.yaml` you only need two things. + +**1. Point the PDS at your metrics backend.** Set the variables above in +`/pds/pds.env`, with `OTEL_EXPORTER_OTLP_METRICS_ENDPOINT` pointed at your own +OTLP endpoint. If your Prometheus runs elsewhere, enable its OTLP receiver +(`--web.enable-otlp-receiver`) and use +`http://your-prometheus:9090/api/v1/otlp/v1/metrics`. If you run an +OpenTelemetry Collector, send to that instead. + +**2. Import the dashboard.** In Grafana, *Dashboards → New → Import → Upload +JSON file*, and pick `dashboards/pds-overview.json`. It expects a Prometheus +data source; select yours when prompted. + +For host metrics, you presumably already run node_exporter. If not, you can +start just that one service from this directory: + +```bash +docker compose up --detach node-exporter +``` + +and add a scrape job to your own Prometheus: + +```yaml +scrape_configs: + - job_name: pds-node + static_configs: + - targets: ['your-pds-host:9100'] +``` + +Note that `compose.yaml` binds node_exporter to `127.0.0.1`, so a Prometheus on +another machine cannot reach it as-is. Either scrape it over an SSH tunnel or a +private network, or change `--web.listen-address`, but if you make it listen on +a public interface, firewall it. `node_exporter` has no authentication. + +## Retention and disk + +Prometheus keeps 15 days by default (`--storage.tsdb.retention.time` in +`compose.yaml`) and stores data in a Docker volume, not under `/pds`. For a +single PDS this is a small amount of data, but it is not +included in a `/pds` backup, and that it does share the host's disk with your +repos and blobs. + +## If a panel is empty + +Metric names come from the OpenTelemetry instrumentation and are translated by +Prometheus's OTLP receiver (dots become underscores, and type and unit suffixes +are appended — `account.created` becomes `account_created_total`). Names can +shift as the instrumentation libraries are upgraded. + +To see what your PDS is actually reporting: + +```bash +curl -s localhost:9090/api/v1/label/__name__/values | tr ',' '\n' | grep -Ei 'account|session|oauth|http_server|nodejs|v8js' +``` + +If a name differs from what a panel queries, edit the panel's query, or open an +issue so the dashboard can be fixed. \ No newline at end of file diff --git a/monitoring/compose.yaml b/monitoring/compose.yaml new file mode 100644 index 0000000..04d5314 --- /dev/null +++ b/monitoring/compose.yaml @@ -0,0 +1,75 @@ +name: pds-monitoring + +# Self-contained monitoring stack for a self-hosted PDS. +# +# docker compose up --detach +# +# All three services use host networking (matching the PDS stack) and bind to +# 127.0.0.1, so nothing here is reachable from the internet. Reach Grafana with +# an SSH tunnel: +# +# ssh -L 3001:localhost:3001 you@your-pds-host +# +# Grafana is on 3001 because the PDS itself owns port 3000. +# +# If you already run Prometheus and Grafana elsewhere, don't use this file -- +# see README.md for the two pieces you need instead. + +services: + node-exporter: + container_name: pds-node-exporter + image: prom/node-exporter:v1 + network_mode: host + pid: host + restart: unless-stopped + command: + - '--path.procfs=/host/proc' + - '--path.sysfs=/host/sys' + - '--path.rootfs=/rootfs' + - '--collector.filesystem.mount-points-exclude=^/(dev|proc|sys|run|var/lib/docker/.+)($$|/)' + - '--web.listen-address=127.0.0.1:9100' + volumes: + - /proc:/host/proc:ro + - /sys:/host/sys:ro + - /:/rootfs:ro + + prometheus: + container_name: pds-prometheus + image: prom/prometheus:v3 + network_mode: host + restart: unless-stopped + command: + - '--config.file=/etc/prometheus/prometheus.yml' + - '--storage.tsdb.path=/prometheus' + - '--storage.tsdb.retention.time=15d' + - '--web.listen-address=127.0.0.1:9090' + # Lets the PDS push OTLP metrics directly to Prometheus, so this stack + # does not need a separate OpenTelemetry Collector. + - '--web.enable-otlp-receiver' + volumes: + - ./prometheus.yml:/etc/prometheus/prometheus.yml:ro + - prometheus-data:/prometheus + + grafana: + container_name: pds-grafana + # Grafana publishes no floating major tag, so this is pinned to an exact + # version and needs a manual bump to upgrade. + image: grafana/grafana:13.1.1 + network_mode: host + depends_on: + - prometheus + restart: unless-stopped + environment: + GF_SERVER_HTTP_ADDR: 127.0.0.1 + GF_SERVER_HTTP_PORT: 3001 + GF_ANALYTICS_REPORTING_ENABLED: false + GF_ANALYTICS_CHECK_FOR_UPDATES: false + GF_USERS_ALLOW_SIGN_UP: false + volumes: + - ./grafana/provisioning:/etc/grafana/provisioning:ro + - ./dashboards:/var/lib/grafana/dashboards:ro + - grafana-data:/var/lib/grafana + +volumes: + prometheus-data: + grafana-data: diff --git a/monitoring/dashboards/pds-overview.json b/monitoring/dashboards/pds-overview.json new file mode 100644 index 0000000..65d39e0 --- /dev/null +++ b/monitoring/dashboards/pds-overview.json @@ -0,0 +1,773 @@ +{ + "uid": "pds-overview", + "title": "PDS Overview", + "description": "Host and PDS metrics for a self-hosted AT Protocol PDS. Host metrics come from node_exporter; PDS metrics are pushed over OTLP.", + "tags": ["pds", "atproto"], + "editable": true, + "schemaVersion": 39, + "version": 1, + "refresh": "1m", + "time": { "from": "now-6h", "to": "now" }, + "timezone": "browser", + "graphTooltip": 1, + "templating": { + "list": [ + { + "name": "instance", + "label": "Host", + "type": "query", + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "query": "label_values(node_uname_info, instance)", + "refresh": 1, + "includeAll": true, + "multi": false, + "current": { "text": "All", "value": "$__all" } + } + ] + }, + "panels": [ + { + "type": "row", + "id": 100, + "title": "PDS — activity", + "gridPos": { "h": 1, "w": 24, "x": 0, "y": 0 }, + "collapsed": false, + "panels": [] + }, + { + "type": "stat", + "id": 1, + "title": "Accounts created", + "description": "Total over the selected time range.", + "gridPos": { "h": 4, "w": 6, "x": 0, "y": 1 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(increase(account_created_total[$__range]))", + "instant": true + } + ], + "options": { + "colorMode": "none", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "short", + "decimals": 0, + "color": { "mode": "fixed", "fixedColor": "text" }, + "noValue": "0" + }, + "overrides": [] + } + }, + { + "type": "stat", + "id": 2, + "title": "Sessions created", + "description": "Total over the selected time range.", + "gridPos": { "h": 4, "w": 6, "x": 6, "y": 1 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(increase(session_created_total[$__range]))", + "instant": true + } + ], + "options": { + "colorMode": "none", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "short", + "decimals": 0, + "color": { "mode": "fixed", "fixedColor": "text" }, + "noValue": "0" + }, + "overrides": [] + } + }, + { + "type": "stat", + "id": 3, + "title": "OAuth authorizations", + "description": "Apps granted access, over the selected time range.", + "gridPos": { "h": 4, "w": 6, "x": 12, "y": 1 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(increase(oauth_authorization_total[$__range]))", + "instant": true + } + ], + "options": { + "colorMode": "none", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "short", + "decimals": 0, + "color": { "mode": "fixed", "fixedColor": "text" }, + "noValue": "0" + }, + "overrides": [] + } + }, + { + "type": "stat", + "id": 4, + "title": "XRPC request rate", + "gridPos": { "h": 4, "w": 6, "x": 18, "y": 1 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(rate(http_server_request_duration_seconds_count[$__rate_interval]))" + } + ], + "options": { + "colorMode": "none", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "reqps", + "decimals": 1, + "color": { "mode": "fixed", "fixedColor": "text" }, + "noValue": "0" + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 5, + "title": "Top XRPC methods by request rate", + "description": "Requires per-route labels on HTTP server metrics — see monitoring/README.md if this panel is empty.", + "gridPos": { "h": 8, "w": 12, "x": 0, "y": 5 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "topk(10, sum by (http_route) (rate(http_server_request_duration_seconds_count[$__rate_interval])))", + "legendFormat": "{{http_route}}" + } + ], + "options": { + "legend": { "displayMode": "table", "placement": "right", "showLegend": true, "calcs": ["mean", "max"] }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "reqps", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 6, + "title": "XRPC request latency", + "gridPos": { "h": 8, "w": 12, "x": 12, "y": 5 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "histogram_quantile(0.50, sum by (le) (rate(http_server_request_duration_seconds_bucket[$__rate_interval])))", + "legendFormat": "p50" + }, + { + "refId": "B", + "expr": "histogram_quantile(0.95, sum by (le) (rate(http_server_request_duration_seconds_bucket[$__rate_interval])))", + "legendFormat": "p95" + }, + { + "refId": "C", + "expr": "histogram_quantile(0.99, sum by (le) (rate(http_server_request_duration_seconds_bucket[$__rate_interval])))", + "legendFormat": "p99" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "s", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 7, + "title": "Account and session activity", + "gridPos": { "h": 8, "w": 12, "x": 0, "y": 13 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum by (source) (rate(session_created_total[$__rate_interval]))", + "legendFormat": "sessions — {{source}}" + }, + { + "refId": "B", + "expr": "sum(rate(account_created_total[$__rate_interval]))", + "legendFormat": "accounts created" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "cps", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 8, + "title": "HTTP responses by status code", + "gridPos": { "h": 8, "w": 12, "x": 12, "y": 13 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum by (http_response_status_code) (rate(http_server_request_duration_seconds_count[$__rate_interval]))", + "legendFormat": "{{http_response_status_code}}" + } + ], + "options": { + "legend": { "displayMode": "table", "placement": "right", "showLegend": true, "calcs": ["mean"] }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "reqps", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 10, + "stacking": { "mode": "normal", "group": "A" }, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "row", + "id": 101, + "title": "PDS — Node.js runtime", + "gridPos": { "h": 1, "w": 24, "x": 0, "y": 21 }, + "collapsed": false, + "panels": [] + }, + { + "type": "timeseries", + "id": 9, + "title": "Event loop utilization", + "description": "Fraction of time the event loop was busy. Sustained values near 1 mean the PDS is CPU-bound.", + "gridPos": { "h": 7, "w": 8, "x": 0, "y": 22 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "nodejs_eventloop_utilization_ratio", + "legendFormat": "utilization" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": false }, + "tooltip": { "mode": "single", "sort": "none" } + }, + "fieldConfig": { + "defaults": { + "unit": "percentunit", + "min": 0, + "max": 1, + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 10, + "showPoints": "never", + "spanNulls": true + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 10, + "title": "Event loop delay", + "gridPos": { "h": 7, "w": 8, "x": 8, "y": 22 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "nodejs_eventloop_delay_p50_seconds", + "legendFormat": "p50" + }, + { + "refId": "B", + "expr": "nodejs_eventloop_delay_p99_seconds", + "legendFormat": "p99" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "s", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 11, + "title": "V8 heap used", + "gridPos": { "h": 7, "w": 8, "x": 16, "y": 22 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(v8js_memory_heap_used_bytes)", + "legendFormat": "heap used" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": false }, + "tooltip": { "mode": "single", "sort": "none" } + }, + "fieldConfig": { + "defaults": { + "unit": "bytes", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 10, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "row", + "id": 102, + "title": "Host", + "gridPos": { "h": 1, "w": 24, "x": 0, "y": 29 }, + "collapsed": false, + "panels": [] + }, + { + "type": "stat", + "id": 12, + "title": "CPU used", + "gridPos": { "h": 4, "w": 6, "x": 0, "y": 30 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "100 - (avg(rate(node_cpu_seconds_total{mode=\"idle\",instance=~\"$instance\"}[$__rate_interval])) * 100)" + } + ], + "options": { + "colorMode": "value", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "percent", + "decimals": 1, + "min": 0, + "max": 100, + "color": { "mode": "thresholds" }, + "thresholds": { + "mode": "absolute", + "steps": [ + { "color": "green", "value": null }, + { "color": "yellow", "value": 75 }, + { "color": "red", "value": 90 } + ] + } + }, + "overrides": [] + } + }, + { + "type": "stat", + "id": 13, + "title": "Memory used", + "gridPos": { "h": 4, "w": 6, "x": 6, "y": 30 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "(1 - (sum(node_memory_MemAvailable_bytes{instance=~\"$instance\"}) / sum(node_memory_MemTotal_bytes{instance=~\"$instance\"}))) * 100" + } + ], + "options": { + "colorMode": "value", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "percent", + "decimals": 1, + "min": 0, + "max": 100, + "color": { "mode": "thresholds" }, + "thresholds": { + "mode": "absolute", + "steps": [ + { "color": "green", "value": null }, + { "color": "yellow", "value": 80 }, + { "color": "red", "value": 92 } + ] + } + }, + "overrides": [] + } + }, + { + "type": "stat", + "id": 14, + "title": "Disk used (/)", + "gridPos": { "h": 4, "w": 6, "x": 12, "y": 30 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "100 - (sum(node_filesystem_avail_bytes{mountpoint=\"/\",instance=~\"$instance\"}) / sum(node_filesystem_size_bytes{mountpoint=\"/\",instance=~\"$instance\"}) * 100)" + } + ], + "options": { + "colorMode": "value", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "percent", + "decimals": 1, + "min": 0, + "max": 100, + "color": { "mode": "thresholds" }, + "thresholds": { + "mode": "absolute", + "steps": [ + { "color": "green", "value": null }, + { "color": "yellow", "value": 75 }, + { "color": "red", "value": 90 } + ] + } + }, + "overrides": [] + } + }, + { + "type": "stat", + "id": 15, + "title": "Load (1m)", + "gridPos": { "h": 4, "w": 6, "x": 18, "y": 30 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(node_load1{instance=~\"$instance\"})" + } + ], + "options": { + "colorMode": "none", + "graphMode": "area", + "textMode": "auto", + "justifyMode": "auto", + "reduceOptions": { "calcs": ["lastNotNull"], "fields": "", "values": false } + }, + "fieldConfig": { + "defaults": { + "unit": "short", + "decimals": 2, + "color": { "mode": "fixed", "fixedColor": "text" } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 16, + "title": "CPU by mode", + "gridPos": { "h": 8, "w": 12, "x": 0, "y": 34 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum by (mode) (rate(node_cpu_seconds_total{mode!=\"idle\",instance=~\"$instance\"}[$__rate_interval]))", + "legendFormat": "{{mode}}" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "short", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 10, + "stacking": { "mode": "normal", "group": "A" }, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0, + "axisLabel": "cores" + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 17, + "title": "Memory", + "gridPos": { "h": 8, "w": 12, "x": 12, "y": 34 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum(node_memory_MemTotal_bytes{instance=~\"$instance\"}) - sum(node_memory_MemAvailable_bytes{instance=~\"$instance\"})", + "legendFormat": "used" + }, + { + "refId": "B", + "expr": "sum(node_memory_MemTotal_bytes{instance=~\"$instance\"})", + "legendFormat": "total" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "bytes", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [ + { + "matcher": { "id": "byName", "options": "total" }, + "properties": [ + { "id": "custom.lineStyle", "value": { "fill": "dash", "dash": [8, 4] } }, + { "id": "color", "value": { "mode": "fixed", "fixedColor": "text" } } + ] + } + ] + } + }, + { + "type": "timeseries", + "id": 18, + "title": "Disk I/O", + "gridPos": { "h": 8, "w": 12, "x": 0, "y": 42 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum by (device) (rate(node_disk_read_bytes_total{device=~\"(sd|vd|nvme|xvd).*\",instance=~\"$instance\"}[$__rate_interval]))", + "legendFormat": "read — {{device}}" + }, + { + "refId": "B", + "expr": "sum by (device) (rate(node_disk_written_bytes_total{device=~\"(sd|vd|nvme|xvd).*\",instance=~\"$instance\"}[$__rate_interval]))", + "legendFormat": "write — {{device}}" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "Bps", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 19, + "title": "Network", + "gridPos": { "h": 8, "w": 12, "x": 12, "y": 42 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "sum by (device) (rate(node_network_receive_bytes_total{device!=\"lo\",instance=~\"$instance\"}[$__rate_interval]))", + "legendFormat": "in — {{device}}" + }, + { + "refId": "B", + "expr": "sum by (device) (rate(node_network_transmit_bytes_total{device!=\"lo\",instance=~\"$instance\"}[$__rate_interval]))", + "legendFormat": "out — {{device}}" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "Bps", + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "axisSoftMin": 0 + } + }, + "overrides": [] + } + }, + { + "type": "timeseries", + "id": 20, + "title": "Filesystem used", + "description": "Watch this one — a PDS that fills its disk stops accepting writes.", + "gridPos": { "h": 7, "w": 24, "x": 0, "y": 50 }, + "datasource": { "type": "prometheus", "uid": "pds-prometheus" }, + "targets": [ + { + "refId": "A", + "expr": "100 - (node_filesystem_avail_bytes{fstype!~\"tmpfs|overlay|squashfs|ramfs\",instance=~\"$instance\"} / node_filesystem_size_bytes{fstype!~\"tmpfs|overlay|squashfs|ramfs\",instance=~\"$instance\"} * 100)", + "legendFormat": "{{mountpoint}}" + } + ], + "options": { + "legend": { "displayMode": "list", "placement": "bottom", "showLegend": true }, + "tooltip": { "mode": "multi", "sort": "desc" } + }, + "fieldConfig": { + "defaults": { + "unit": "percent", + "min": 0, + "max": 100, + "color": { "mode": "palette-classic" }, + "custom": { + "drawStyle": "line", + "lineWidth": 2, + "fillOpacity": 0, + "showPoints": "never", + "spanNulls": true, + "thresholdsStyle": { "mode": "dashed" } + }, + "thresholds": { + "mode": "absolute", + "steps": [ + { "color": "transparent", "value": null }, + { "color": "red", "value": 90 } + ] + } + }, + "overrides": [] + } + } + ] +} diff --git a/monitoring/grafana/provisioning/alerting/.gitkeep b/monitoring/grafana/provisioning/alerting/.gitkeep new file mode 100644 index 0000000..2932741 --- /dev/null +++ b/monitoring/grafana/provisioning/alerting/.gitkeep @@ -0,0 +1 @@ +# Grafana logs an error at startup if this directory is missing. diff --git a/monitoring/grafana/provisioning/dashboards/pds.yml b/monitoring/grafana/provisioning/dashboards/pds.yml new file mode 100644 index 0000000..55059bd --- /dev/null +++ b/monitoring/grafana/provisioning/dashboards/pds.yml @@ -0,0 +1,11 @@ +apiVersion: 1 + +providers: + - name: PDS + type: file + folder: PDS + updateIntervalSeconds: 30 + allowUiUpdates: false + options: + path: /var/lib/grafana/dashboards + foldersFromFilesStructure: false diff --git a/monitoring/grafana/provisioning/datasources/prometheus.yml b/monitoring/grafana/provisioning/datasources/prometheus.yml new file mode 100644 index 0000000..82e8b82 --- /dev/null +++ b/monitoring/grafana/provisioning/datasources/prometheus.yml @@ -0,0 +1,12 @@ +apiVersion: 1 + +datasources: + - name: Prometheus + uid: pds-prometheus + type: prometheus + access: proxy + url: http://127.0.0.1:9090 + isDefault: true + editable: false + jsonData: + timeInterval: 15s diff --git a/monitoring/grafana/provisioning/plugins/.gitkeep b/monitoring/grafana/provisioning/plugins/.gitkeep new file mode 100644 index 0000000..2932741 --- /dev/null +++ b/monitoring/grafana/provisioning/plugins/.gitkeep @@ -0,0 +1 @@ +# Grafana logs an error at startup if this directory is missing. diff --git a/monitoring/prometheus.yml b/monitoring/prometheus.yml new file mode 100644 index 0000000..1087300 --- /dev/null +++ b/monitoring/prometheus.yml @@ -0,0 +1,28 @@ +global: + scrape_interval: 15s + evaluation_interval: 15s + +# Host metrics are scraped. PDS metrics are *pushed* by the PDS over OTLP to +# /api/v1/otlp/v1/metrics (enabled by --web.enable-otlp-receiver in +# compose.yaml), so the PDS needs no scrape job here. +scrape_configs: + - job_name: node + static_configs: + - targets: ['127.0.0.1:9100'] + + - job_name: prometheus + static_configs: + - targets: ['127.0.0.1:9090'] + +otlp: + # Keep OTel resource attributes that identify which PDS a series came from. + promote_resource_attributes: + - service.name + - service.version + - service.instance.id + - host.name + +storage: + tsdb: + # OTLP is a push protocol, so samples can arrive slightly out of order. + out_of_order_time_window: 30m