Hub Monitoring Dashboard
The Hub exports operational metrics to Google Cloud Monitoring through OpenTelemetry. Scion includes a ready-made Cloud Monitoring dashboard for these metrics at deploy/monitoring/dashboards/scion-hub.json. The dashboard shows each Hub replica as its own line, which makes it most useful in HA hosted deployments with several replicas. It works the same way for a single-node Hub.
Use the dashboard for rates and per-instance history. The Hub’s admin Health page shows only the current state of the instance that served the request.
What the dashboard shows
Section titled “What the dashboard shows”| Section | Charts | Metrics |
|---|---|---|
| Database connection pool | Active, idle and max connections and waits per second, per pool (store, events) |
scion.db.pool.connections.* |
| Broker dispatch | Claimed, done and failed per second; stuck pending messages; intent-to-done latency (p50, p95, p99) | scion.dispatch.* |
| Event notifications | Publish-to-deliver lag (p50, p95, p99); drops per second by reason | scion.db.notify.* |
| Launch reaper | Ticks per second by outcome; row errors per second; time disarmed | scion.launch_reaper.* |
Every chart is grouped by the service_instance_id label, so each Hub replica gets its own line. The dashboard also has a scion_hub_id filter at the top. All replicas of one Hub share a Hub ID, so this filter selects one Hub deployment when several Hubs export to the same project.
Prerequisites
Section titled “Prerequisites”- Hub metrics export is on. The Hub exports metrics when
server.hub.gcp_project_id(SCION_SERVER_HUB_GCPPROJECTID) is set. The startup log then containsHub OTel metrics export enabled. The metric groups can be turned off one by one withSCION_METRICS_DB_POOL,SCION_METRICS_DB_NOTIFYandSCION_METRICS_DISPATCH. A disabled group shows as empty charts. - The Hub can write metrics. The Hub’s service account needs
roles/monitoring.metricWriteron that project. - You can create dashboards. Importing needs
roles/monitoring.dashboardEditor(or an equivalent role) on the project.
The notification charts need the Postgres database driver. A Hub on SQLite does not record the scion.db.notify.* metrics.
Import the dashboard
Section titled “Import the dashboard”The JSON file contains no project ID. Choose the project when you import it:
gcloud monitoring dashboards create \ --project=PROJECT_ID \ --config-from-file=deploy/monitoring/dashboards/scion-hub.jsonReplace PROJECT_ID with the project in server.hub.gcp_project_id. The command prints the new dashboard’s resource name (projects/…/dashboards/DASHBOARD_ID). Open it in the Google Cloud console under Monitoring > Dashboards > Scion Hub.
To find the dashboard later:
gcloud monitoring dashboards list --project=PROJECT_ID \ --filter='displayName="Scion Hub"' --format='value(name)'To pick up a newer version of the file, delete the old dashboard (gcloud monitoring dashboards delete DASHBOARD_ID) and create it again. You can also paste the file into the console’s JSON editor for an existing dashboard. Charts update within one export interval (60 seconds) of new data arriving.
How Hub metrics appear in Cloud Monitoring
Section titled “How Hub metrics appear in Cloud Monitoring”The Hub uses the Google Cloud OpenTelemetry metric exporter. If you build your own charts or alert policies, use these mappings:
-
Metric type:
workload.googleapis.com/followed by the OpenTelemetry name, dots included. For example,scion.dispatch.claimedbecomesworkload.googleapis.com/scion.dispatch.claimed. -
Replica label: each Hub process sets the
service.instance.idresource attribute to its own instance ID, a random UUID created at startup (newInstanceIDinpkg/hub/server.go). It is prefixed with the pod name only when thePOD_NAMEenvironment variable is set. The UUID part is always there, so no two Hub processes share an instance ID: not replicas that share a Hub ID, and not a restarted container that keeps its pod name. Nothing in the configuration sets or overrides it. Hub traces (withSCION_TRACING_ENABLED) carry the sameservice.instance.id, so a replica’s spans match its metric series. The Helm chart setsPOD_NAMEfrom the pod’s name, so chart legends read as<pod name>-<UUID>; other deployments show bare UUIDs unless they set it. The instance ID becomes the metric labelservice_instance_id. You don’t configure it. It changes every time a replica restarts, so each restart or rollout starts a new set of series and the old ones stop receiving points. -
Deployment labels: the
scion.hub.idresource attribute becomes the metric labelscion_hub_id. It comes from the Hub ID (server.hub.hub_id, environment variableSCION_SERVER_HUB_HUBID), which every replica of an HA Hub shares. Use it to filter by deployment, not to tell replicas apart. Thescion.hub.nameresource attribute becomesscion_hub_name. Don’t group by it: when no name is configured it falls back to the host name, which is not a stable replica identity. -
Label names: the exporter replaces every character that is not a letter or digit with
_. Point attributes such asoutcome(reaper ticks),reason(notification drops) andpool(connection pools) are metric labels too. -
Pool label: the Hub has two connection pools, and every
scion.db.pool.*point carries apoollabel naming one:store(the main database pool) orevents(the Postgres event pool). Group or filter pool charts by it so the two pools are not summed. -
Monitored resource:
generic_task, withjobset toscion-hubandtask_idset to the replica’s instance ID.locationisglobalandnamespaceis empty. -
Kinds:
OpenTelemetry instrument Metric kind Value type Typical aligner Counter (for example scion.dispatch.done)CUMULATIVEINT64ALIGN_RATEGauge (for example scion.db.pool.connections.active)GAUGEINT64orDOUBLEALIGN_MEANorALIGN_MAXHistogram (for example scion.dispatch.intent_to_done.duration, in ms)CUMULATIVEDISTRIBUTIONALIGN_DELTAwith a percentile reducer
Link from the Health page
Section titled “Link from the Health page”A link to this dashboard from the Hub’s admin Health page is planned.
Related guides
Section titled “Related guides”- Metrics & OpenTelemetry: agent and Hub telemetry configuration
- Observability: logs, traces and troubleshooting
- HA Overview: running several Hub replicas