Keyboard shortcuts

Press ← or → to navigate between chapters

Press S or / to search in the book

Press ? to show this help

Press Esc to hide this help

A production deployment

This page takes you from nothing to a first federated query over two openEHR CDRs you already run. It walks the steps in order and links the reference page for each, so read those for every key and every edge case. The specification leaves deployment to each federation, so no specification governs this page: our own design.

Every value below is a placeholder: hosts under example.org, identifiers in the urn:oid:2.999 example arc. The two members are cdr-a and cdr-b, and the gateway’s public base URL is https://gateway.example.org/fed.

scripts/checks/production-guide.sh assembles every TOML block on this page into one ferrofed.toml and one registry.toml, runs ferrofed config check over them, and serves them behind the shipped reverse proxy. Each block opens with a comment naming the file it belongs in.

1. What you need

WhatWhyReference
Two openEHR CDRs that serve ITS-REST 1.1.0, ad hoc AQL includedthe members; each answers AQL scoped to one ehr_idWhat FerroFED runs beside
An identity provider for your callersevery caller presents an RFC 9068 access token; step 4 configures KeycloakClient authentication
A PIX Managerresolves each patient to the ehr_id each member holdsIdentity resolution
An Audit Record Repository with the FHIR Feed of ITI-20takes the audit and access recordsThe audit trail
A Linux host with Docker Compose, or a Kubernetes clusterruns the published imageThe container image
A DNS name and a TLS certificate for the gateway, and nginx or another reverse proxythe public address clients, nodes and identity services reachstep 8
Prometheus, optionalscrapes the admin listenerMetrics

Decide three values before you start, because several settings repeat them:

  • the public base URL, here https://gateway.example.org/fed: its path is server.base_path, and server.public_url names it once, from which auth.audience, the JWK Set’s URL and the PMIR feed’s URL follow (The public base URL);
  • the federation id, federation.id, which OPTIONS {base}/ reports;
  • each member’s node id and endpoint id, which the registry, the PIX Manager’s domains and the onward credentials all key on.

Prepare each CDR

Agree each of these with the member’s operator before you admit the member. The first five are the conditions of membership (§12b.1, §12b.2, N42a); the rest are what the gateway needs to reach the member.

  • A unique system_id. The CDR stamps one openEHR system_id into every EHR and version it creates, and no other member uses it (§12b.2). Find it by asking the CDR itself, at the member, as the member’s operator:

    SELECT DISTINCT e/system_id/value FROM EHR e
    

    Expect one value. If the CDR reports several, ask its operator which one it stamps into new EHRs, and map any other system whose versions it holds with a [[creating_system]] entry (The registry).

  • ehr_ids are version-4 UUIDs, or come from a scheme with the same collision resistance (§12b.2).

  • No ehr_id is reused, across restores, migrations and test-data resets (§12b.2). The evidence is the member’s documented procedures.

  • No foreign ehr_id is adopted. An imported EHR gets a fresh ehr_id (§12b.2).

  • Each ehr_id reaches the PIX Manager in the member’s ehr_id domain, with the patient identifier your clients use, from the moment the EHR is created, and the EHRs that exist today in a bulk load (§5.5; step 6).

  • ITS-REST is reachable from the gateway at one base URL per endpoint, over https whenever the gateway sends it a credential (What must travel encrypted).

  • The member accepts the gateway’s credential: a bearer token, a user and a password, or an OAuth 2.0 client registered under the gateway’s client_id with the gateway’s JWK Set location (step 5).

  • The member makes its own access decision and enforces consent (§13.2, N26, N27). If it marks a consent refusal with an ITS-REST Error code, write the code down for the registry’s consent_refusal_codes (Consent).

  • The member verifies who asks, recommended: it reads the openEHR-federation-client token the gateway signs, with the gateway’s JWK Set, and records the caller in its own audit trail (Verifying it at the node).

2. Install

The release lane attests every image and every binary. Verify what you pull before you run it, with the GitHub CLI:

gh attestation verify oci://ghcr.io/ferrohealth/ferrofed:X.Y.Z \
  --repo FerroHEALTH/FerroFED \
  --signer-workflow FerroHEALTH/FerroFED/.github/workflows/release-image.yml

Then pin the image by the digest you verified, ghcr.io/ferrohealth/ferrofed@sha256:…; docker pull prints it on its Digest: line. For a binary, verify the tarball against release-build.yml instead (The release binaries).

Pick one layout. Both mount the configuration at /etc/ferrofed/, the secrets at /run/secrets/ferrofed/ and a durable volume at /var/lib/ferrofed/, the paths every block on this page uses:

  • One host: the release’s compose.yaml, ferrofed.toml and registry.toml (The gateway from a release). The gateway’s port is published on 127.0.0.1:8080, where the reverse proxy of step 8 reaches it.
  • Kubernetes: the StatefulSet example in deploy/kubernetes/ (Kubernetes). Run the reverse proxy as a container in the same pod, or put your ingress in front of the Service with TLS on the gateway’s listener.

Edit the release’s example files as you go through the steps below. The gateway itself, and the address the proxy passes to:

# ferrofed.toml
profile = "production"

[server]
listen = "127.0.0.1:8080"
base_path = "/fed"
public_url = "https://gateway.example.org/fed"
trusted_proxies = ["127.0.0.1"]
forwarded_header = "x-forwarded-for"
drain_delay_ms = 5000
shutdown_timeout_ms = 30000

[telemetry]
format = "json"

[federation]
id = "example-federation"
node_selection = "ask-all"
per_node_timeout_ms = 10000
overall_timeout_ms = 25000

The image sets FERROFED__SERVER__LISTEN=0.0.0.0:8080, which overrides listen inside the container; the published port decides who reaches it. trusted_proxies and forwarded_header take the client’s address from the reverse proxy of step 8 on loopback, and from no one else (Behind a reverse proxy). drain_delay_ms keeps the listener open while a load balancer stops routing to a stopping gateway (Stopping without dropping a request).

3. The registry

The registry document names the members: one [[organisation]] per operator, one [[node]] per CDR with its system_id, and one [[endpoint]] per ITS-REST base URL (The registry). Add each endpoint suspended: the gateway sends a suspended endpoint nothing, and the admission check of step 9 still reaches it.

# registry.toml
[[organisation]]
id = "org-a"
name = "Example organisation A"

[[organisation]]
id = "org-b"
name = "Example organisation B"

[[node]]
id = "cdr-a"
organisation = "org-a"
system_id = "cdr-a.example.org"

[[node]]
id = "cdr-b"
organisation = "org-b"
system_id = "cdr-b.example.org"

[[endpoint]]
id = "cdr-a-query"
node = "cdr-a"
url = "https://cdr-a.example.org/openehr"
connection_type = "openehr-rest-query"
managing_organisation = "org-a"
status = "suspended"

[[endpoint]]
id = "cdr-b-query"
node = "cdr-b"
url = "https://cdr-b.example.org/openehr"
connection_type = "openehr-rest-query"
managing_organisation = "org-b"
status = "suspended"
# ferrofed.toml
[registry]
document = "/etc/ferrofed/registry.toml"
format = "toml"

The registry refuses a second member with the system_id of another, and reloads on SIGHUP without a restart (Reloading the registry).

4. Client authentication

Every request to {base}/v1/ and OPTIONS {base}/ carries an access token the gateway verifies before it reads anything else (Client authentication). The token must be an RFC 9068 access token:

  • its JOSE header names the type at+jwt;
  • it carries iss, exp, aud, sub, client_id, iat and jti, and aud names auth.audience, which is server.public_url unless you set it;
  • its scope holds the SMART on openEHR scopes of each operation, such as user/aql-*.s for a federated query (Scopes per route);
  • it declares a purpose of use, as extensions.ihe_iua.purpose_of_use, an array of FHIR Coding, or as RFC 9396 authorization_details (Purpose of use).

An issuer recipe: Keycloak

Keycloak 26.2 added a client setting that types its access tokens at+jwt. The recipe below was run on 2026-10-06 against Keycloak 26.8.0: the gateway admitted a user’s token and a client-credentials token made this way, and answered a federated query for each. With the at+jwt setting off, it refused the same user’s token as “not typed at+jwt”, and without the professional mapper it refused the client-credentials token as naming no natural person.

Run it with Keycloak’s admin CLI, kcadm.sh, against your Keycloak’s public address. It creates a realm, one client scope per SMART on openEHR scope, a clinical application that signs users in, and a reporting service that uses the client-credentials grant. The realm’s acr.loa.map names the levels of Keycloak’s authentication flows, so its tokens carry acr as urn:example:loa:substantial for level 1 and urn:example:loa:high for level 2, the values the gateway’s [auth.issuer.assurance] reads below. Keycloak 26.8.0 wrote level 1 into a token after a password sign-in and into a client-credentials token. A real realm reaches each level only by a flow whose means meets it, which Regulation (EU) No 910/2014 Art 8(2) defines and your identity provider’s assessment establishes; the password sign-in here does not:

kcadm.sh config credentials --server https://idp.example.org --realm master --user admin
kcadm.sh create realms -s realm=ferrofed -s enabled=true \
  -s 'attributes."acr.loa.map"="{\"urn:example:loa:substantial\":1,\"urn:example:loa:high\":2}"'
for scope in 'user/aql-*.s' 'user/composition-*.r' 'system/aql-*.s'; do
  kcadm.sh create client-scopes -r ferrofed -s "name=$scope" -s protocol=openid-connect \
    -s 'attributes."include.in.token.scope"=true' \
    -s 'attributes."display.on.consent.screen"=false'
done
kcadm.sh create clients -r ferrofed -s clientId=example-clinical-app \
  -s publicClient=false -s standardFlowEnabled=true \
  -s 'redirectUris=["https://app.example.org/callback"]' \
  -s 'attributes."access.token.header.type.rfc9068"=true'
kcadm.sh create clients -r ferrofed -s clientId=example-reporting-service \
  -s publicClient=false -s standardFlowEnabled=false -s serviceAccountsEnabled=true \
  -s 'attributes."access.token.header.type.rfc9068"=true'

Each client then gets its scopes as default client scopes, loses the profile and email scopes, whose personal data the gateway never reads, and gets its protocol mappers: the audience and the purpose of use for both, client_id for the clinical application, and the professional the reporting service acts for:

id_of() { kcadm.sh get clients -r ferrofed -q "clientId=$1" --fields id --format csv --noquotes; }
scope_of() {
  kcadm.sh get client-scopes -r ferrofed --fields id,name --format csv --noquotes |
    awk -F, -v n="$1" '$2 == n { print $1 }'
}
app="$(id_of example-clinical-app)"
svc="$(id_of example-reporting-service)"
for scope in 'user/aql-*.s' 'user/composition-*.r'; do
  kcadm.sh update "clients/$app/default-client-scopes/$(scope_of "$scope")" -r ferrofed
done
kcadm.sh update "clients/$svc/default-client-scopes/$(scope_of 'system/aql-*.s')" -r ferrofed
for client in "$app" "$svc"; do
  for scope in profile email; do
    kcadm.sh delete "clients/$client/default-client-scopes/$(scope_of "$scope")" -r ferrofed
  done
  kcadm.sh create "clients/$client/protocol-mappers/models" -r ferrofed -f audience.json
  kcadm.sh create "clients/$client/protocol-mappers/models" -r ferrofed -f purpose-of-use.json
done
kcadm.sh create "clients/$app/protocol-mappers/models" -r ferrofed -f client-id.json
kcadm.sh create "clients/$svc/protocol-mappers/models" -r ferrofed -f professional.json

audience.json puts the gateway’s audience in aud:

{
  "name": "ferrofed-audience",
  "protocol": "openid-connect",
  "protocolMapper": "oidc-audience-mapper",
  "config": {
    "included.custom.audience": "https://gateway.example.org/fed",
    "access.token.claim": "true",
    "id.token.claim": "false",
    "introspection.token.claim": "true"
  }
}

purpose-of-use.json declares the purpose of use, here treatment (TREAT of HL7 v3 ActReason), in the IHE IUA extension. A hardcoded claim gives one purpose per client, so a client that requests data for another purpose gets a client of its own:

{
  "name": "purpose-of-use",
  "protocol": "openid-connect",
  "protocolMapper": "oidc-hardcoded-claim-mapper",
  "config": {
    "claim.name": "extensions.ihe_iua.purpose_of_use",
    "claim.value": "[{\"system\": \"http://terminology.hl7.org/CodeSystem/v3-ActReason\", \"code\": \"TREAT\"}]",
    "jsonType.label": "JSON",
    "access.token.claim": "true",
    "id.token.claim": "false",
    "userinfo.token.claim": "false",
    "introspection.token.claim": "true"
  }
}

client-id.json adds client_id to the user’s token. RFC 9068 requires the claim, and Keycloak 26.8.0 wrote it by itself into the service account’s token and not into the user’s:

{
  "name": "client-id",
  "protocol": "openid-connect",
  "protocolMapper": "oidc-hardcoded-claim-mapper",
  "config": {
    "claim.name": "client_id",
    "claim.value": "example-clinical-app",
    "jsonType.label": "String",
    "access.token.claim": "true",
    "id.token.claim": "false",
    "userinfo.token.claim": "false",
    "introspection.token.claim": "true"
  }
}

professional.json names the professional the reporting service acts for, in the IHE IUA extension. A system/ scope acts without a user, so the gateway admits the service to patient data only for a professional its token names. The value is a synthetic identifier; yours is the identifier the professional’s national authority issued:

{
  "name": "professional",
  "protocol": "openid-connect",
  "protocolMapper": "oidc-hardcoded-claim-mapper",
  "config": {
    "claim.name": "extensions.ihe_iua.national_provider_identifier",
    "claim.value": "urn:oid:2.999.7.1.42",
    "jsonType.label": "String",
    "access.token.claim": "true",
    "id.token.claim": "false",
    "userinfo.token.claim": "false",
    "introspection.token.claim": "true"
  }
}

A user who signs in to example-clinical-app then gets a token whose scope reads user/composition-*.r user/aql-*.s, and the reporting service a token with system/aql-*.s that names its professional, both with acr urn:example:loa:substantial. Keycloak’s realm issuer is https://idp.example.org/realms/ferrofed; copy issuer and jwks_uri from https://idp.example.org/realms/ferrofed/.well-known/openid-configuration, since the gateway compares iss exactly:

# ferrofed.toml
[auth]
clock_skew_s = 60

[[auth.issuer]]
issuer = "https://idp.example.org/realms/ferrofed"
jwks_uri = "https://idp.example.org/realms/ferrofed/protocol/openid-connect/certs"
backend_clients = ["example-reporting-service"]
client_tokens_act_for_professional = true

[auth.issuer.assurance]
claim = "acr"
minimum = "substantial"
substantial = ["urn:example:loa:substantial"]
high = ["urn:example:loa:high"]

system/aql-* counts only for a client listed in backend_clients. client_tokens_act_for_professional admits the reporting service to patient data for the professional professional.json names; a client token of this issuer that names none is refused 401 natural-person-required. [auth.issuer.assurance] reads the realm’s acr and refuses a token below substantial with 401 authentication-assurance-insufficient (Professionals and assurance). An operator who uses the operator console needs the issuer’s operator_scope as well (The operator surface).

5. Onward credentials

The gateway never forwards a caller’s token to a node. It authenticates to each endpoint with that endpoint’s own credential, one section per endpoint id, and signs the caller’s identity onto every request with its own key (Onward credentials). Here cdr-a takes a bearer token and cdr-b registers the gateway as an OAuth 2.0 client, authenticated by an assertion the signing key signs:

# ferrofed.toml
[credentials."cdr-a-query"]
bearer_token_file = "/run/secrets/ferrofed/cdr-a-token"

[credentials."cdr-b-query".oauth2]
grant = "client_credentials"
client_auth = "private_key_jwt"
token_endpoint = "https://auth.cdr-b.example.org/oauth2/token"
client_id = "ferrofed-gateway"
scope = "system/aql-*.s system/composition-*.r"

[signing]
key_file = "/run/secrets/ferrofed/signing-key.pem"

Make the signing key, and give every secret file to the gateway’s user alone:

openssl genpkey -algorithm EC -pkeyopt ec_paramgen_curve:P-384 -out secrets/signing-key.pem
sudo chown -R 65532:65532 secrets && sudo chmod 0400 secrets/*

Use ec_paramgen_curve:P-256 instead when a member’s authorization server holds to the FAPI 2.0 Security Profile (Signing keys and the JWK Set). The JWK Set is served at https://gateway.example.org/fed/.well-known/jwks.json, on the public address of step 8, and signing.jwks_uri takes that value from server.public_url; give it to cdr-b’s operator with the client_id, and to every member that verifies the caller token.

6. Identity resolution

A patient query names its patient by an identifier and its namespace. The gateway asks the PIX Manager for the patient’s ehr_id in each member’s ehr_id domain (ITI-83), and sends each member a query scoped to that ehr_id alone (§5.2, N3). Without a cross-reference every patient query is a 424 (Choosing one).

The Manager answers only from what its feeds delivered, so it must hold, for every patient, the identifier clients name and each member’s ehr_id in that member’s domain. Each member, or the integration engine beside it, feeds its own domain: ITI-104 or the PMIR ITI-93 feed when it creates an EHR, and once, a bulk load of the EHRs it already holds (Keeping the PIX Manager current). Register each domain at the Manager before you add the member.

SanteMPI is the PIX Manager the end-to-end lane verifies against two members, and the block below is its shape: an OAuth 2.0 client-credentials grant with the secret in the request body (A PIX Manager verified with FerroFED: SanteMPI).

# ferrofed.toml
[[pixm.manager]]
url = "https://mpi.example.org/fhir/"

[pixm.manager.members]
"cdr-a" = "urn:oid:2.999.10"
"cdr-b" = "urn:oid:2.999.20"

[pixm.manager.credentials.oauth2]
grant = "client_credentials"
token_endpoint = "https://mpi.example.org/auth/oauth2_token"
client_id = "ferrofed-pix-consumer"
client_auth = "client_secret_post"
client_secret_file = "/run/secrets/ferrofed/mpi-consumer-secret"
scope = "*"

Every member of the registry is resolved by exactly one Manager, so a member missing from [pixm.manager.members] refuses the configuration. When your hospital MPI does not answer ITI-83, run a Manager for the federation beside it, or let PDQm translate a local identifier first (When the hospital MPI does not speak PIXm).

To hear of identity merges as they happen, subscribe at a PMIR Patient Identity Registry. It sends each change to the feed route under the base, which step 8 opens to the Registry alone (The identity feed):

# ferrofed.toml
[pmir]
url = "https://pmir.example.org/fhir"
feed_token_file = "/run/secrets/ferrofed/pmir-feed-token"

The Registry sends the feed to pmir.callback_url, which is https://gateway.example.org/fed/pmir/feed here, the feed route under server.public_url.

7. The access log

Every IHE transaction the gateway makes is audited, and every access to patient data it intermediates is recorded with the caller it verified (The audit trail, #623). Both go to your Audit Record Repository as FHIR AuditEvents, from a spool on the durable volume:

# ferrofed.toml
[audit]
destination = "repository"

[audit.repository]
url = "https://arr.example.org/fhir"
hostname = "gateway.example.org"
spool_dir = "/var/lib/ferrofed/audit-feed-spool"

Outside the development profile, config check refuses this configuration without destination, naming audit.destination. With the access log of #623, a gateway with a registry refuses log as well, because the log target names no caller and no patient.

A record counts as recorded once it is on the spool’s disk. Keep the volume on encrypted storage that outlives the container, one spool per replica (Losing a spool).

8. TLS and the public address

The listener speaks plain HTTP by default. Choose one of two shapes:

  • A reverse proxy beside the gateway terminates TLS and reaches the gateway over loopback, on the same host or in the same pod. This page ships one for nginx.
  • TLS on the listener, with [server.tls], where the hop from your proxy or load balancer to the gateway crosses a network others share. Add client_ca_file to admit the proxy alone (TLS on the listeners).

ferrofed.conf is the nginx configuration for the settings above, and every release carries it with its checksum. Download it from the release you run, check it, set your server name and certificate, and reload nginx:

for f in ferrofed.conf ferrofed.conf.sha256sum; do
  curl -LO "https://github.com/FerroHEALTH/FerroFED/releases/download/vX.Y.Z/$f"
done
sha256sum -c ferrofed.conf.sha256sum && sudo cp ferrofed.conf /etc/nginx/conf.d/
nginx -t && nginx -s reload

It passes these routes, each under the base:

RouteWho reaches it
GET and OPTIONS {base}/your clients
{base}/v1/ and belowyour clients; the gateway authenticates each one
{base}/operator/ and belowthe operator console
GET {base}/.well-known/jwks.jsonevery member’s authorization server, and every member that verifies the caller token
POST {base}/pmir/feedthe PMIR Patient Identity Registry alone; name its address in the allow line
{base}/health and belowstopped at the proxy with 403: the dependency report names every member endpoint, so probe the gateway on its own port
every other path404 at the proxy

It also:

  • passes each request URI as the client sent it, so a version uid’s :: and every encoded character reach the gateway unchanged;
  • keeps its own timeouts above the gateway’s 30-second request timeout and its body limit above the gateway’s 1 MiB, so a client gets the gateway’s own answer and error code;
  • sets Strict-Transport-Security, which the gateway does not send;
  • sets X-Forwarded-For to the address the request came from, replacing any a client sent, and drops Forwarded, so the gateway, which trusts this proxy alone, names the client’s address and never one a client wrote;
  • writes an access log line with the gateway’s request id in place of the request line, because the GET form of a query and a read by subject carry the patient identifier in the query string. The request id finds the gateway’s own log line.

scripts/checks/production-guide.sh runs the file in front of a gateway with this page’s configuration, and checks each route in the table, the header and the log line.

Under the Dutch Nuts grant, the gateway also serves its DID document at the path its did:web names, outside the base; pass that path as well (The Nuts grant).

9. Check, start and admit

Check the configuration before every start, with the same image you run. It resolves the configuration exactly as serve would, secret files included, opens no store, binds no socket and writes nothing, so it needs no durable volume: the audit spool directory is checked where it stands and created by serve (Running it):

docker run --rm -v "$PWD/ferrofed.toml:/etc/ferrofed/ferrofed.toml:ro" \
  -v "$PWD/registry.toml:/etc/ferrofed/registry.toml:ro" \
  -v "$PWD/secrets:/run/secrets/ferrofed:ro" \
  ghcr.io/ferrohealth/ferrofed@sha256:… config check --config /etc/ferrofed/ferrofed.toml

A refused file exits 78 with one line naming the key at fault, never a secret (Running it). Then start the gateway and ask for its readiness on the host:

docker compose up --wait
curl http://127.0.0.1:8080/fed/health/readiness
curl -H "Authorization: Bearer $OPERATOR_TOKEN" \
  http://127.0.0.1:8080/fed/operator/dependencies

Readiness is 200 once the gateway serves; no member and no identity service gates it. The dependency report shows the last state the gateway observed of each member, the resolver and the audit repository to a token with the operator scope (The dependency report).

Admit each member before it serves (§12b.1, N42a). The admission check creates synthetic EHRs at the member and checks each condition a test can reach, so agree the run with the member’s operator first (Admitting a node):

docker compose exec ferrofed /usr/local/bin/ferrofed admission check --endpoint cdr-a-query

When the member’s governance forbids test data in its production CDR, run the full check against a staging copy of it, and check the production CDR with --read-only, which writes nothing and names every condition it leaves unproven (A run without writes).

When both members pass, remove status = "suspended" from their endpoints and reload the registry:

docker compose kill -s SIGHUP ferrofed

The Kubernetes example has no reload path into the pod yet (#636): change the ConfigMap and restart the StatefulSet’s pods.

10. The first federated query

Get a token from the issuer of step 4. The reporting service’s client-credentials token needs no browser:

token="$(curl -s https://idp.example.org/realms/ferrofed/protocol/openid-connect/token \
  -d grant_type=client_credentials -d client_id=example-reporting-service \
  -d client_secret="$REPORTING_SERVICE_SECRET" | jq -r .access_token)"

Then send one ordinary ITS-REST query for a patient both members know:

curl -s https://gateway.example.org/fed/v1/query/aql \
  -H "Authorization: Bearer $token" \
  -H 'Content-Type: application/json' -d @- <<'EOF' | jq .meta.federation
{"q": "SELECT c/uid/value FROM EHR e CONTAINS COMPOSITION c WHERE e/ehr_status/subject/external_ref/id/value = 'example-0001' AND e/ehr_status/subject/external_ref/namespace = 'urn:oid:2.999.1'"}
EOF

The answer is one ITS-REST RESULT_SET, and meta.federation says how it was made (What a client gets back):

  • complete is true when every member in scope answered.
  • endpoints has one entry per member. active with a row_count means the member was asked and answered. not-resolved means the PIX Manager holds no ehr_id for the patient there: the member was not asked, and complete is false. node-error, time-out and offline name a member that failed, with the reason in error.
  • The openEHR-federation-endpoint and openEHR-federation-system-id response headers name the members that contributed rows.

A 401 names its reason in the WWW-Authenticate challenge, a 403 in its error code, and a 424 means a member that may hold the patient could not be resolved (Errors and status codes).

11. Operations

Turn on the admin listener, and scrape GET /metrics there. It stays on loopback unless you allow otherwise, and never sits under the base path. Off loopback, the scrape needs a token, which your Prometheus sends with authorization.credentials_file, and a write action such as the stored-query distribution needs an access token with your issuer’s operator_scope (Metrics):

# ferrofed.toml
[metrics]
listen = "127.0.0.1:9464"
# off loopback: listen = "0.0.0.0:9464", allow_remote = true, and
# scrape_token_file = "/run/secrets/ferrofed/metrics-scrape-token"
  • Alerts. Load the shipped dashboard and alert rules, and add an alert on the scrape job’s up (Dashboard and alert rules).
  • Logs. Ship the JSON log to a store with access control and a retention period. It carries no patient identifier; an outbound-gate-stopped event is an incident (Hardening).
  • Restarts. On SIGTERM readiness turns 503, the listener stays open for drain_delay_ms, the requests in flight get shutdown_timeout_ms, and the bindings get bindings_drain_timeout_ms to stop their processes. Keep the runtime’s grace period above all three (Stopping without dropping a request).
  • Several replicas. The replicas share nothing in memory. A follow-up write needs the client to name its endpoint, or the balancer to keep a client on one replica; the stored-query registry needs PostgreSQL; each replica keeps its own audit spool (Running several replicas).
  • Upgrades. Read the release’s upgrade notes, keep a copy of the configuration and the image digest you run, verify the new image as in step 2, and run its config check over your files as in step 9 before you replace the image (Upgrading). Going back has its own steps for the configuration, the stored-query store and the audit spools (Rollback).
  • Signing-key rotation. Rotate in three rolling restarts (Rotating the signing key).
  • Hardening. Work through the hardening checklist before the gateway reaches real patient data.