Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

Metrics

The gateway counts what it already observes: the requests its clients send it, by route and status, and how long they take; the requests it sends to each member node, and to its resolver, localizer and demographics service, and how long they take; the integrity incidents it raises; the security events it logs; the requests its limits refuse; and the registry reloads. One OpenTelemetry meter provider holds the counts. A Prometheus server scrapes them from GET /metrics on an admin listener of their own, and the gateway can also push them to an OpenTelemetry collector over OTLP. Both surfaces read the same provider, so a metric never exists on one and not the other. Both are off by default.

The dependency report, GET {base}/operator/dependencies, stays as it is: it reports the last state the gateway observed of each member, and the metrics count every request over time.

Turning it on

[metrics]
listen = "127.0.0.1:9464"     # the admin listener; unset, nothing listens
allow_remote = false          # true lets listen name a non-loopback address
scrape_token_file = "/run/secrets/ferrofed/metrics-scrape-token"   # the bearer token a scrape must carry; unset, the scrape is open
otlp_endpoint = "http://127.0.0.1:4317"   # an OTLP gRPC collector; unset, nothing is pushed

The admin listener serves GET /metrics and the operator’s stored-query distribution (POST /admin/stored-queries/{name}/{version}/distribute), answers every other path 404, and never sits under the base path. It is never the gateway’s own listener, so no client of the federation reaches it.

Who the admin listener serves

RequestAdmittedRefused
GET /metrics, no scrape_token setevery peernever
GET /metrics, scrape_token seta scrape with Authorization: Bearer <the scrape token>401 unauthenticated with a WWW-Authenticate: Bearer challenge, for a scrape with no token, another token or another scheme
A write action, such as the distributiona caller whose access token carries the operator_scope of its [[auth.issuer]], from any peer401 unauthenticated without a token that verifies, 403 scope-insufficient without the operator scope
A write action under profile = "development"as above, and a loopback peer (127.0.0.0/8, ::1) that sends no credentialas above

A write action is verified by the same client authentication as the gateway’s own listener, against the issuers of [auth] and in its mode, and needs no purpose of use. An issuer that names no operator_scope admits no operator, so with none the write actions are refused to every caller outside the development profile. The scrape token admits the scrape alone, never a write action. Every refusal is counted under ferrofed_security_events_total as admin-write-refused or scrape-refused, and logged under the target ferrofed::security with its reason, never the credential.

Every admitted write action is recorded under ferrofed::security in two lines. The first, event admin-write-admitted, is written before the action has any effect and is counted under the same name. The second, event admin-write-finished, is written when the action ends, with its outcome: the status it was answered, or abandoned when it never answered because its client went away or its handler failed. Each line carries the operator’s issuer and subject, the method and route template as action, and the time as at (RFC 3339), and nothing else: no token, header, path value, body or stored query. A loopback peer the development profile admitted without a credential names no issuer or subject. Neither the issuer nor the subject is ever a metric label.

The subject is the one the issuer put in the operator’s token, and it may name a person. The gateway keeps it as issued, because a keyed hash under a key that changes at every start could not attribute an action after a restart. Keep these lines as long as you keep your other security records, and restrict who reads them. No [telemetry] filter quiets the ferrofed::security target below info, so the records are written whatever the filter says.

scrape_token, or scrape_token_file with the token in a file read at start, sets the scrape token; give one or the other. Prometheus sends it with authorization in the scrape job:

scrape_configs:
  - job_name: ferrofed
    authorization:
      type: Bearer
      credentials_file: /etc/prometheus/secrets/ferrofed-scrape-token
    static_configs:
      - targets: ["gateway.example.org:9464"]

Mutual TLS authenticates the scrape as well: with [metrics.tls] and its client_ca_file, every client of the listener must present a certificate that CA signed (TLS on the listeners). Give the scraper such a certificate with tls_config in its job.

config check says who a listener that is not on a loopback address answers. serve and config check refuse:

  • a listen address that is not a loopback address, such as 0.0.0.0:9464, unless allow_remote = true is set. Set it only when the address is reachable from your scraper and your operators and nothing else, for example inside a pod network a network policy closes; the policy is defence in depth, and holds only where the cluster’s network plugin enforces it;
  • outside profile = "development", a listen address that is not a loopback address with nothing to authenticate the scrape: neither scrape_token nor [metrics.tls] client_ca_file;
  • scrape_token and scrape_token_file both set, and a scrape_token_file that cannot be read or holds nothing;
  • a listen address equal to server.listen;
  • an otlp_endpoint that is not an http:// URL. The push speaks gRPC without TLS, so run the collector beside the gateway, on the same host or in the same pod, and let it forward over TLS;
  • an otlp_endpoint with a user name or a password in it, outside profile = "development", since that credential would travel in cleartext (What must travel encrypted).

The push sends every 60 seconds; the standard OTEL_METRIC_EXPORT_INTERVAL environment variable, in milliseconds, changes the interval. A push that fails is logged at WARN under the target opentelemetry-otlp, and the counts stay readable on /metrics. On SIGTERM the gateway pushes once more after the drain.

[metrics] takes effect on a restart only: a reload that changes it logs the key as needing a restart, like [server].

The metrics

The gateway names its instruments the OpenTelemetry way, and the Prometheus exporter renders each name with _ for ., _total after a counter, and the unit after a histogram.

Prometheus nameInstrumentTypeLabelsCounts
ferrofed_integrity_incidents_totalferrofed.integrity.incidentscounterkindthe integrity incidents the gateway emitted, one per incident line under ferrofed::integrity
ferrofed_node_requests_totalferrofed.node.requestscounterendpoint, outcomethe requests the gateway sent to a member endpoint
ferrofed_node_request_duration_secondsferrofed.node.request.duration (unit s)histogramendpointthe time a member endpoint took to answer, with the series _bucket (and le), _sum and _count
ferrofed_consent_prefilter_requests_totalferrofed.consent.prefilter.requestscounteroutcome, reasonthe calls to the consent pre-filter, by denied, no-signal, not-asked, unavailable or partial; a not-asked call, one that never reached the consent service, also carries a reason: namespace (the patient is named in a namespace the service is not asked by, such as a pseudonymised BSN for Mitz) caller-claims (the caller’s token does not state the claims the service is asked on behalf of), caller-claims-invalid (the token states them in a form the service’s question does not take) or patient-value (the patient’s value is not one the identifier the service is asked by takes)
ferrofed_localizer_requests_totalferrofed.localizer.requestscounteroutcomethe calls to the localizer, by candidates, no-records, not-configured, unavailable, or audit-failed for an XCPD exchange whose audit message could not be recorded
ferrofed_demographics_requests_totalferrofed.demographics.requestscounteroutcomethe calls to the demographics step, by identified, no-match, ambiguous, unavailable, or audit-failed for an exchange whose audit record could not be stored
ferrofed_http_requests_totalferrofed.http.requestscounterhttp_request_method, http_route, status_classthe requests the gateway answered, by the route template and the status class 1xx to 5xx, a refused one included
http_server_request_duration_secondshttp.server.request.duration (unit s)histogramhttp_request_method, http_route, http_response_status_code, url_scheme, and error_type on a 5xxthe time from receiving a request to answering it, the OpenTelemetry HTTP server metric (https://opentelemetry.io/docs/specs/semconv/http/http-metrics/)
http_server_active_requestshttp.server.active_requestsgaugehttp_request_method, url_schemethe requests being served now
ferrofed_resolver_requests_totalferrofed.resolver.requestscounteroutcomethe calls to the cross-reference resolver, a PIX Manager or the development cross-reference, one per patient lookup across the members asked, by resolved, not-resolved, unavailable (the service failed for a member) or time-out (its budget ran out)
ferrofed_resolver_request_duration_secondsferrofed.resolver.request.duration (unit s)histogramnonethe time each resolver call took
ferrofed_localizer_request_duration_secondsferrofed.localizer.request.duration (unit s)histogramnonethe time each localizer call took, an XCPD or NVI exchange; a not-configured call asks nothing and is not timed
ferrofed_demographics_request_duration_secondsferrofed.demographics.request.duration (unit s)histogramnonethe time each PDQm call took
ferrofed_security_events_totalferrofed.security.eventscounterevent, and reason on caller-refusedthe security events of the log targets ferrofed::security: a caller refused at client authentication by its reason, an issuer’s key set or introspection endpoint that cannot be had, and every identifier-hygiene event: a query refused before dispatch, a patient predicate stripped, a request the outbound gate stopped, a query parameter or a declared value refused, a probe refused, a stored-query definition refused, a patient grant’s confinement, an admin listener write action or scrape refused at the listener’s authentication, and an admin listener write action an admitted operator ran
ferrofed_overload_refusals_totalferrofed.overload.refusalscounterlimit, and endpoint on node-in-flightthe requests a limit refused (Overload protection)
ferrofed_registry_reloads_totalferrofed.registry.reloadscounterresultthe registry reloads SIGHUP asked for
ferrofed_identity_feed_messages_totalferrofed.identity_feed.messagescounterresultthe ITI-93 messages the identity feed received, by applied, refused (not held to the PMIR profiles), unauthenticated (no feed token) or audit-failed (its audit record could not be stored, and nothing was applied)
ferrofed_audit_spool_eventsferrofed.audit.spool.eventsgaugenonethe ITI-20 audit messages waiting in the spools for the audit repository and the FHIR Feed repository of [audit], summed; present only with one configured
ferrofed_audit_spool_bytesferrofed.audit.spool.bytesgaugenonethe bytes the spool holds, its quarantine included
ferrofed_audit_quarantinedferrofed.audit.quarantinedgaugenonethe audit messages in the spool’s quarantine, which could not be read or were no whole frame
ferrofed_audit_delivered_totalferrofed.audit.deliveredcounternonethe ITI-20 audit messages delivered to the audit repository since the process started
ferrofed_audit_retries_totalferrofed.audit.retriescounternonethe failed attempts to deliver, each followed by a backoff
ferrofed_audit_refused_totalferrofed.audit.refusedcounternonethe audit messages the spool refused for want of room under spool_max_events or spool_max_bytes, the messages queued for a write counted with those stored; each fails its exchange as an audit failure
target_infothe resourcegaugeservice_name, service_version, telemetry_sdk_*always 1: the gateway and its version

Every label value comes from a closed set or from your registry document, never from a request, so no patient identifier, query text, header value or path reaches the surface:

LabelValues
kindEhrIdCollision, IndexInsertCollision, LearnedCreatingSystemConflict, RegisteredCreatingSystemConflict
endpointan endpoint id of the registry document
outcomeactive, node-error, time-out, offline, consent-denied
resultapplied, refused
http_routea route template of the gateway, such as /v1/query/aql or /v1/ehr/{ehr_id}/composition, under the base path, or <unmatched> for a path no route names; a path identifier is written as its parameter name, never its value
http_request_methodGET, POST, PUT, DELETE, OPTIONS, HEAD, PATCH, CONNECT, TRACE, or _OTHER for any other method
http_response_status_code, error_typethe HTTP status the gateway answered
status_class1xx, 2xx, 3xx, 4xx, 5xx
url_schemehttp where the listener speaks plain HTTP and TLS ends in front of it, https where it serves [server.tls] itself
eventcaller-refused, key-set-unavailable, introspection-unavailable, aql-refused, patient-predicate-stripped, subject-parameters-consumed, outbound-gate-stopped, query-parameter-refused, parameter-value-refused, ehr-id-probe-refused, definition-subject-literal, held-definition-refused, patient-confinement, patient-context-unavailable, admin-write-refused, scrape-refused, admin-write-admitted: the event field of the log line
reasonon ferrofed_security_events_total, the reason the WWW-Authenticate challenge names: missing, malformed, algorithm, type, issuer, key, signature, expired, not-yet-valid, audience, inactive, unavailable, operation, scope, demographic-client, purpose-of-use, patient-context, patient-demographic, natural-person, assurance, contact-point-attributes, correlation
limitconcurrency, caller-rate, node-in-flight
lea bucket bound in seconds: on the member, resolver, localizer and demographics histograms 0.005, 0.01, 0.025, 0.05, 0.1, 0.25, 0.5, 1, 2.5, 5, 10, 30, +Inf; on http_server_request_duration_seconds the bounds the OpenTelemetry HTTP conventions advise, 0.005, 0.01, 0.025, 0.05, 0.075, 0.1, 0.25, 0.5, 0.75, 1, 2.5, 5, 7.5, 10, with 30 added for the request timeout, and +Inf

The incident, reload and security event counters, and the concurrency and caller-rate refusals, show every label value at 0 from the start, so an alert on their increase works from the first scrape. A node request series appears with the first request to that endpoint, and an endpoint a reload removes keeps its series until a restart.

Which calls the node request series cover

Six kinds of call send a request to a member. Each request one of them sends is counted once in ferrofed_node_requests_total and timed once in ferrofed_node_request_duration_seconds, under the same endpoint, so the histogram’s _count equals the counter summed over outcome. A request that never left the gateway is in neither: a member that was not-resolved, excluded or not-localized, and a request the gateway could not send, for want of a client or a credential, or because the identifier-hygiene gate withheld it. §11.1 has no status for a request the gateway could not send, so a member record in meta.federation still reports that member offline; the series do not count it. A request whose deadline passed before it left is not counted either, in any of the six calls, because the node was never asked. The client still sees the budget run out: a time-out in the member record, or a 504 (node-timeout) for a routed request or a probe. A call the gateway fails with a 500 because one of its own tasks panicked counts no member at all, since a defect in the gateway is never a node’s time-out.

CallRequests countedoutcome read fromTime recorded
A federated query, GET or POST {base}/v1/query/aql, and a stored-query invocationone per member sent the queryits §11.1 status in meta.federationits latency_ms
A fan-out template uploadone per member sent the uploadits §11.1 status in meta.federationits latency_ms
A stored-query distribution, a PUT naming members or the admin listener’s repairone per member sent the definitionits §11.1 status in meta.federationits latency_ms
A stored-query drift check, a GET naming membersone per member asked for its copyits §11.1 status in meta.federationits latency_ms
A request routed to one node: the {base}/v1/ehr/ area, GET {base}/v1/ehr?subject_id=…, a definition request naming one endpoint, a demographic requestonethe node’s answerfrom sending it to the answer
The ask-all probe, GET /ehr/{ehr_id} at every member for a read whose owner no earlier step namedone per member probedthe node’s answerfrom sending it to the answer, or to the moment the overall budget ran out

The per-member record a federated query, a template upload, a distribution and a drift check write is the one the client reads in meta.federation, and the counter follows it exactly. A routed request and a probe have no such record, so their outcome is read from the node’s answer by the same rules. Each outcome therefore covers these calls:

outcomeCovers
activea member that answered with success; for a routed request or a probe, any answer below 500, a 404 included, so a probe that finds no EHR at a member is active
node-errorin a member record, any answer that is not a success, a 4xx included (except a consent refusal the registry names), and a drift check whose copy differs from the registry’s definition or is missing; for a routed request or a probe, a 5xx answer; in every call, a node that refused the gateway’s onward credentials
time-outa member that gave no answer before its per-node deadline, and a member still being waited on when the overall budget ran out
offlinea member the gateway sent a request to and could not reach
consent-denieda request whose node answered 403 with a consent refusal code the registry lists for it (Consent): a federated query member, a routed or by-subject read, or an ask-all probe, counted so whether or not the client’s answer withholds it (Withholding consent exclusions); a member a consent pre-filter dropped is sent no request and is not counted

Dashboard and alert rules

Every release attaches a Grafana dashboard, ferrofed-dashboard.json, and a Prometheus rule file, ferrofed-alerts.yaml; both are in the repository under deploy/observability/. Import the dashboard and pick your Prometheus as its datasource: it charts the requests by status class and route, the members by outcome and answer time, the resolver, localizer and demographics calls, the security events, the limits’ refusals, the integrity incidents and the audit spool. Load the rule file through rule_files in prometheus.yml, or wrap its groups in a PrometheusRule for the Prometheus Operator. Its alerts carry severity: page or severity: ticket:

AlertFires when
FerroFEDServerErrorsover 5% of the answers are 5xx for 10 minutes
FerroFEDSlowAnswersthe 95th percentile answer takes over 20 seconds for 10 minutes
FerroFEDOverloadedthe concurrency limit refuses requests for 5 minutes
FerroFEDCallerRateLimiteda caller is rate limited for 15 minutes
FerroFEDMemberFailinga member endpoint fails over 10% of its requests for 10 minutes
FerroFEDMemberCapSaturateda member’s in-flight cap stays full for 5 minutes
FerroFEDResolverFailing, FerroFEDLocalizerFailingthe resolver or the localizer does not answer for 5 minutes
FerroFEDIntegrityIncidentan integrity incident is raised
FerroFEDRegistryReloadRefuseda registry reload is refused
FerroFEDAuditSpoolBacklog, FerroFEDAuditRefusedthe audit spool holds over 1000 records for 15 minutes, or refuses one
FerroFEDOutboundGateStoppedthe outbound gate stops a request that would have carried a patient identifier
FerroFEDCallerRefusalscallers are refused at authentication over once a second for 10 minutes, by reason
FerroFEDAuthenticationUnavailablean issuer’s key set or introspection endpoint cannot be had

The thresholds are a starting point; tune them to your traffic. Add an alert on your scrape job’s up as well, since a gateway that does not answer the scrape counts nothing. scripts/checks/observability.sh holds both files to the metrics the gateway exports, and runs promtool check rules.

The Kubernetes example (deploy/kubernetes/) serves the admin listener on port 9464 of each pod, with the scrape token in the metrics-scrape-token key of the ferrofed-secrets Secret, annotates the pods for a Prometheus that reads prometheus.io/scrape, and opens the port to the Prometheus pods of the monitoring namespace alone with a network policy. Give the Prometheus job the same token with authorization.credentials_file. The gateway authenticates the scrape and every write action by itself, so the policy is defence in depth.

The gateway does not call a webhook. An incident is counted, and its log line under ferrofed::integrity carries the routing ids you act on; route the alert from your Prometheus or your collector.