Cluster API (MAIN ↔ LB)¶
A load balancer used to reach MAIN's MariaDB and Redis directly: its crons, daemons and
streaming endpoints held the panel's database credentials and wrote MAIN's tables. The
cluster API replaces that with one signed, encrypted HTTP channel the node opens to MAIN.
This page is the protocol and the operator's view of it; the design record with every
decision and its history is docs/adr/0004-cluster-api.md,
and the code seams each call goes through are in Cluster seams.
Plain HTTP is the default
The channel does not need TLS. Every request is signed and its body encrypted with keys
derived from the node's token, so a panel with no certificate loses nothing by talking
over http://. HTTPS is optional and, when a certificate verifies, preferred — see
Transport.
What travels¶
The node always dials MAIN; MAIN never connects to a node (its instructions wait for the
node to ask for them). Every call is POST /cluster/v1/<op> on MAIN's HTTP broadcast port,
or on cluster_api_port when one is set.
| Direction | Carried as | Examples |
|---|---|---|
| Node → MAIN, needs an answer | an op on the control lane | hello, heartbeat, token_refresh, conn_admit, commands, ack |
| Node → MAIN, a report | an op on an ingest lane | events, streams, conn_snapshot, queue_claim, queue_update, queue_enqueue, config, recording_complete, vod_analysis, artefact |
| MAIN → node | a signed command the node's next commands call collects |
node.rpc, node.root, node.cache, conn.kill_worker, conn.drop, conn.close, config.changed, artefact.fetch |
Two ops the plan lists are deliberately never served: rpc_result (a command's result comes
back inline with its ack, which takes 64 KiB) and stream_bundle (the replica's streams
section carries a stream's whole record, and a cache miss reads it from disk). They stay in
the op list because nginx's ingest lane is rendered from it, and the API refuses an op it
does not serve.
Lanes¶
Events are spooled by priority and drained by the agent:
| Lane | What | Pace |
|---|---|---|
| P0 | a viewer opening or closing, a stream's state | at once |
| P1 | logs, the node's inventory, its audit | batched |
| P2 | HLS touches | batched, only while MAIN takes them |
MAIN serves the control ops on the cluster_ctl FPM pool and the ingest ops on
cluster_ingest, sized by cluster_ingest_concurrency; half its permits are reserved for
P0 so a flood of logs can never delay a viewer. A node that arrives while MAIN is starting
is refused with STARTING and a back-off, never a 500.
Authentication¶
There is no shared secret in a file and no bearer credential on the wire.
- Enrolment. MAIN gives the node an identity (
node_uuid, an Ed25519 key pair the node generates itself, MAIN's pinned public key) over SSH at install, or through a one-time code the operator reads out (cluster:enrol-code, approved withcluster:enrol-approve). - The token is minted inside the
xcvm_coreextension from the panel's root secret and the licence, bound to the node's uuid, generation and epoch. It never exists in PHP as plaintext: the node receives it sealed to its own key. - Every request carries a timestamp, a nonce and an HMAC over a canonical representation of the request, with the body encrypted (AES-256-GCM). MAIN replies the same way. A replayed request is refused by the nonce store; a request whose clock is off by more than the allowed skew is refused.
- Rotation. The token's lifetime is
lb_token_rotation_min(5–1440, default 60). The agent refreshes at half its life; a refused token re-keys rather than stopping the node. Revoking the licence stops new tokens being minted at all, which is how a fleet is fenced.
Transport and endpoints¶
cluster_transport:
| Value | Meaning |
|---|---|
auto (default) |
HTTPS first when MAIN's own certificate verifies, else plain HTTP |
http |
plain HTTP only |
https_preferred |
HTTPS first, plain HTTP as the fallback |
https_required |
HTTPS only; refused unless MAIN's certificate verifies and every active node has already reached MAIN over HTTPS (its agent reports the https feature) |
The URLs a node dials, their order and the transport are MAIN's policy, versioned by
cluster_policy_ver. A node adopts a policy that is not older than the one it holds, and
remembers the plain-HTTP URLs of every policy it has seen: under https_required with a
broken certificate, the signed challenge over plain HTTP is its way back.
Changing MAIN's address, its HTTPS port or cluster_api_port keeps the old endpoint served
for seven days so no node is lost; cluster:endpoint list shows what is kept and
cron:cluster releases a port once every node has moved off it.
The policy also carries the fleet's heartbeat (lb_telemetry_interval_sec, 1–3 s), so
changing it reaches every node with the next policy and no agent restart.
Flows and modes¶
What a node does through the API is switched per node, on the Cluster Nodes page. A flow moves one part of its work to the agent and stands the legacy path down.
| Bit | Flow | Moves |
|---|---|---|
| 1 | TELEMETRY | the server's stats: MAIN writes the servers row from the heartbeats |
| 2 | COMMANDS | kills, cache jobs and root actions arrive as signed commands |
| 4 | LOGS | client, stream, error, activity and on-demand records become log.* events |
| 8 | STREAMS | the node's streams come from its replica, its runtime state from its own store |
| 16 | CONTENT | recordings, VOD analysis and the encoding queue go through MAIN |
| 32 | CONFIG | the node's caches are built from the replica, not MAIN's database |
| 64 | CONNECTIONS | the node's viewers live in its agent's registry |
| 128 | DATAPLANE | relays and other servers' files go through the agent's loopback proxy, with tickets instead of the stream secret |
mode is the node's independence:
| Mode | Name | Meaning |
|---|---|---|
| 0 | legacy | reads MAIN's database as it always did |
| 1 | hybrid | boots from its replica, may still reach MAIN's database |
| 2 | api | every connect to MAIN's MariaDB or Redis is refused in code (LbDatabaseAccessException) |
Moving a node up is gated (ClusterAdmin::modeGate()): mode 1 needs CONFIG; mode 2 needs
every flow, the data plane included, and the node's own connect audit clean — zero MySQL and
zero Redis connects for seven days. Moving down is always allowed, because it is the way
back. A crontab row whose role is legacy (today cron:users) is not sent to a node in
mode 2 at all.
The node replica¶
While CONFIG is on, the node's caches are built from what the agent stored, not from a
query. MAIN serves the sections; cluster:apply turns them into the caches the node's
readers use (a shadow diff before the flow is on, so an operator sees what would change):
| Section | Carries |
|---|---|
settings |
the settings a node reads, never a secret |
servers |
every server's row, without liveness or telemetry |
node |
the node's own row and its node settings |
crontab |
the enabled rows whose role fits the node's mode |
cluster |
MAIN's URLs, the transport, the heartbeat, the panel keys, min_proto |
secrets |
the viewer-token secret and OPENSSL_EXTRA, each with the value it replaced and how long that is still accepted |
bouquets, categories |
as cron:cache builds them |
streams (R2) |
one record per stream the node holds, kept by deltas and a resync; with DATAPLANE on, its relay and file tickets too |
Settings reference¶
| Setting | Bounds | What reads it |
|---|---|---|
cluster_api_enabled |
0/1 | the whole API; needs cluster:init and the extension |
cluster_api_port |
0 or 1024–65535 | 0 serves the API on MAIN's HTTP broadcast port |
cluster_main_host |
hostname | MAIN's DNS name in the URLs nodes dial |
cluster_transport |
see above | the policy |
lb_token_rotation_min |
5–1440 (60) | the token's lifetime; grace is clamp(L/4, 5, 60) |
lb_revocation_mode |
graceful | hard | how fast a fleet stops after a licence revocation |
lb_telemetry_interval_sec |
1–3 (2) | the fleet's heartbeat, carried by the policy |
cluster_offline_after_sec |
10–300 (30) | silence before MAIN marks a node offline |
cluster_orphan_conn_ttl_sec |
30–3600 (120) | silence before MAIN purges a node's viewers |
lb_offline_admission |
local | allow | deny | admitting viewers while MAIN is unreachable |
cluster_kill_on_line_disable |
0/1 (1) | a disabled, locked or expired line loses its sessions |
cluster_ingest_concurrency |
1–64 (6) | MAIN's ingest permits; half reserved for P0 |
lb_new_node_mode |
legacy | api | the mode a newly installed LB enrols at |
servers_stats_retention_days |
1–365 (30) | cron:cleanup prunes servers_stats |
cluster_audit_retention_days |
1–365 (30) | cron:cleanup prunes cluster_audit |
cluster_db_allowlist (+_extra) |
0/1 | firewalls 3306/6379 on MAIN to the fleet |
lb_scan_roots |
paths | the directories the node's scan RPC may list |
lb_partition_tolerance_h, lb_fence_drain_min |
0–24 (12), 0–60 (10) | the lease's window past token expiry, and the drain after it |
lb_lease_fence |
0/1 (0) | a node stops serving when its lease runs out |
Operating it¶
# MAIN, once: create the cluster root and record the panel keys
console.php cluster:init
# Enrol a load balancer over SSH (or at install time, automatically)
console.php server:enrol <serverID>
# Without SSH: a one-time code the operator reads out, then approves by its SAS
console.php cluster:enrol-code <serverID>
console.php cluster:enrol-approve <serverID> <SAS>
# The node's own side, run by its agent after every change
console.php cluster:apply # --from-disk at boot
console.php cluster:exec # one signed command
# MAIN's endpoint and pools
console.php cluster:nginx # render and reload the API's nginx config
console.php cluster:pools # start or resize the API's FPM pools
console.php cluster:endpoint list # old ports and URLs still served
console.php cluster:maintain-stats # servers_stats indexes, built online (cron:cleanup starts it)
# Before switching CONNECTIONS on: load the node's viewers into its agent
console.php cluster:seed-connections <serverID>
# Every node (MAIN too, any mode) keeps the agent, the fanout daemon and xcvm_core
# current from GitHub itself, hourly from cron:root_signals; by hand, as root:
console.php fanout_binary [fanout|agent] [force]
console.php xcvm_core [force]
# Firewall MariaDB and Redis to the fleet (check first)
console.php cluster:db-allowlist status | apply | undo
# Phase 9: a mode-2 node gives up MAIN's credentials (asks first; --wait=<s> waits for the revoke)
console.php cluster:strip-credentials <serverID> [--yes] [--wait=<seconds>]
# MAIN reads other servers' files and relays through its own agent (on | off | rekey | status)
console.php cluster:main-dataplane on
# Rotate the panel's DB password (MAIN), and set it on a node MAIN cannot reach (node, root)
console.php cluster:rotate-db-password [--yes] [--password-stdin]
echo "$NEW_PASSWORD" | console.php cluster:set-db-password
# Disaster recovery of the cluster root
console.php cluster:export-keys /path/bundle
console.php cluster:import-keys /path/bundle
The Cluster Nodes page is where a node is approved, its flows switched, its mode moved and
its state read (epoch, token expiry, the fence window that follows it — token expiry plus
lb_partition_tolerance_h, then lb_fence_drain_min — its queued commands, last seen, agent
version and architecture, settings misses, connect audit). It warns when the licence is
suspended (and, with lb_lease_fence on, by when the fleet stops at the latest) and when
MAIN's certificate expires within 14 days while the nodes may dial HTTPS. It shows the figures
MAIN records: command delivery and ack latency (p50/p99 over the last hour), queue depth, the
ingest permits in use per lane, the cluster_ctl pool's listen queue, and the recent audit.
Rotate all tokens now sends token.rotate_now to every active node that takes commands.
Saving a different activation key on the dashboard does the same, once the extension accepts it.
The Servers list's row menu carries the same per-node actions (mode up/down, rotate,
enrolment code, a link to the node's flows) and shows the cluster:reenrol command to run
for a re-enrolment over SSH. Every decision is written to cluster_audit, which
cron:cluster prunes.
A cutover, in order¶
cluster:initon MAIN,cluster_api_enabledon,cluster:nginx,cluster:pools.- Enrol the node; approve it if it came by code.
- Switch TELEMETRY, then COMMANDS, then LOGS — each is reversible and visible on the page.
- Switch STREAMS and CONTENT; watch the node's settings misses and its audit.
- Switch CONFIG, then
mode_upto 1: the node now boots from its replica. cluster:seed-connections, then CONNECTIONS.- DATAPLANE, once the node's parents and the servers whose files it reads are MAIN or active nodes, and its agent runs the relay proxy (the page refuses the switch otherwise): relays and file reads go through its agent from each stream's next start.
- Leave it for a week. When the node's connect audit shows zero MySQL and zero Redis
connects for seven days,
mode_upto 2. cluster:db-allowlist applyonce every node is in mode 2.- Drop DB credentials on the node (or
cluster:strip-credentials): the node'sconfig.encloses MAIN's DB and Redis credentials, then MAIN revokes its grant. There is no undo from the page: rolling back needs a config with credentials and a new grant.
Limits¶
- The data plane (Phase 8) covers what a node pulls from another server, and only
while its DATAPLANE flow is on. A child relaying a stream, a VOD or subtitle fetch and a
created channel's items go through the agent's loopback proxy (
127.0.0.1:31290), which signs each upstream connect with a panel-signed ticket from the R2 streams section (a relay ticket valid 24 h, a file ticket 6 h, both minted anew every 3 h and never moving a record's ETag or version) and checks each 4 MiB chunk of a file against its owner's signed digest. What it does not cover: - A parent or owner that cannot check a ticket (a legacy server, not enrolled or not active) is still reached with the legacy URL and its password, even with the flow on.
- Streams started before the flow was switched keep their URL until they restart.
- Relay bytes are authenticated at connect only, not framed: over plain HTTP they have no integrity (plan, D11). Files do.
- The secret still appears on the node's own loopback: the local RTMP output
(
rtmp://127.0.0.1/live/<id>?password=) and the recorder's pull from its own/admin/liveand/admin/timeshift. - MAIN pulls through it only once an operator runs
cluster:main-dataplane on: MAIN then gets a data-plane key of its own and an entry in the signed node list, and its source probe, a node's certbot log, and the relays and files of the streams it runs go throughxc_agent run -role main. Off (the default), MAIN keeps the legacy URLs. MAIN's key sits inconfig/cluster/main_agent.json(0600), outsidexcvm_core, as a node's does;cluster:main-dataplane rekeyreplaces it and raises its generation. - A file ticket names the owner's box key as it was when minted: an owner re-enrolled since cannot open it until the next epoch's ticket (at most 3 h).
- The node's own legacy
/apistays served with the flow on.api_legacy.conf(Phase 8's second increment) retires it only once nothing reads the node's files with the legacygetFileURL any more: its own flow on, and every server of the cluster, MAIN included, an active node with its data plane on. MAIN counts oncecluster:main-dataplane onis set. A proxy is never a node, so a cluster with a proxy keeps every node's/api. - The flow can be switched on only for a node whose agent runs the loopback relay proxy
(it says
relayat hello): updatexc_agentfirst. The agent publishes the loopback key (relay.key) only while it holds127.0.0.1:31290, and the node's PHP checks that the listener belongs to the key's owner. While the port is not the agent's (held by another user, or the agent is stopped), the node's relays and file reads fail and are retried: they never fall back to the stream secret. An agent that cannot bind the port tells MAIN in every heartbeat: the Cluster Nodes page shows relay port down with the error, andserver:diagnosenames it. The port is not freed for it: an operator finds the holder on the node. /xfilehas its own rate limit (50 requests/s per server, burst 100, answered with a 429 the agent retries), apart from the viewers' 20 requests/s.cluster:rotate-stream-secretdoes not exist: retiring the password is Phase 9's.- The licence lease is issued, checked, and enforced only behind a switch (Phase 9).
Every token MAIN hands a node (enrolment over SSH or by code,
token_refresh,token_rekey) carries a lease signed byxcvm_core, capped atmin(token_exp + lb_partition_tolerance_h, iat + 26 h); when the extension refuses one (no licence, its clock gate, a revoked generation) the token goes out without it, and a refusal for the licence dispatchesClusterLicenceLapsedEvent.xc_agentverifies each lease it receives (panel signature, its node, server and generation, the window on its estimate of MAIN's time), keeps the newest in its state file and prints it withxc_agent lease. The fence that acts on it is built in the node's PHP (Core\Cluster\NodeLease) behindlb_lease_fence, which is off by default: with it on, past the lease'sexpno new viewer starts on the node (stream/auth.php), and pastlb_fence_drain_minmore the sessions still running stop (segment.php,key.php,live.php), and the agent drops every viewer the fanout serves, judging the same lease, clock and settings itself. It judges the lease state the agent writes toconfig/cluster/lease_state.jsonat every heartbeat interval, whether or not MAIN answers: the lease'sexp, and MAIN's clock carried forward on the node's monotonic clock from MAIN's last authenticated statement, so moving the node's wall clock does not move MAIN's time as the fence reckons it, and restarting the agent resumes it. Every uncertainty serves: the switch off, a legacy node, no file or a file the agent stopped refreshing, no lease, no anchor on MAIN's clock. The switch must be on before a licence lapses — it reaches a node in the replica'ssettingssection, which a panel without a licence cannot sign. Where the node'sxcvm_coreoffers it (cluster_lease_state), the extension judges the lease itself — against its own pin of the panel key and an anchor on MAIN's clock that runs on the monotonic clock and never moves back — and the agent's file is the fallback. The same verdict makeslicense_valid()true on the node, soLicenseGatelets it use fanout. The extension's verdict needs the node'score.pin: the SSH install pins it over its own session, and every other node that takes root commands is pinned bycron:clusterwith a signednode.root pin_core(the core badge on the Cluster Nodes page, or Pin core to send it now). - The credential lockdown is partly built (Phase 9). A node in mode 2 gives up MAIN's
credentials with a signed
node.root strip_db_credentials(or a credential-freenode.root install_config), run byxcvm_coreas root; when its ack reports a config without credentials, MAIN revokes the node's grant (XC_VM::db_revoke) and recordscluster_nodes.db_revoked_at(Domain\Cluster\DbCredentials). Only an operator sends the strip (Drop DB credentials,cluster:strip-credentials), andlb_new_node_mode=apiis still refused (api_mode_allowedis false): the cutover stays the operator's decision. A node MAIN already keeps credential-free (mode 2, or a revoked grant) stays so: a reinstall over SSH packs it a credential-free config, re-enrols it in mode 2 and grants it nothing, and no grant path (Re-authorise MySQL,tools mysql) reaches its host.cluster:rotate-db-passwordrotates the panel's DB password throughXC_VM::db_set_passwordand sends each node below mode 2 that takes root commands a signednode.root rotate_dbwith the new password SEALed to its box key, which its root side opens and hands toXC_VM::config_set_db(onlydb.passchanges). Every other load balancer that still uses its grant keeps the old password until an operator runscluster:set-db-passwordon it. The password never rides a command in the clear.cluster:rotate-credentialsrotates the Redis password the same way, and the manualcluster:lockdownexists too (ADR 0004, Phase 9's fifth and seventh increments). - The viewer-token secret can be replaced gracefully (the value it replaces stays readable for ten minutes, fleet-wide), but a full rotation — re-encrypting what is stored under it — is Phase 9's.
- A node's
whitelist_ipsis the admin's: a node no longer publishes its own addresses, because that column grants the legacy/apiallowlist.