Cluster seams¶
The MAIN ↔ LB API plan (docs/superpowers/specs/2026-09-21-main-lb-api-communication-design.md)
replaces the node's direct access to MAIN's MySQL and Redis with a signed API. Phase 0
gets the code ready without changing any behaviour. Every place a node reaches MAIN's
state, and every place MAIN reaches a node, goes through one class. Each class has one
backend today, the legacy one, which does exactly what the call sites used to do.
Later phases swap that backend and leave the call sites alone.
Each seam has a hook for tests and for the future transport (useSink(),
useLoader() or useTransport()). Passing null restores the legacy backend.
Node → MAIN¶
| Seam | Wraps | Legacy backend | API backend (phase) |
|---|---|---|---|
Core\Cluster\SignalDispatcher |
the 47 INSERT INTO signals sites (kill, cache jobs, root actions) |
LegacySqlSignalSink |
commands and events (4/5) |
Domain\Stream\StreamStateWriter |
a node's runtime state in streams_servers; refuses any column outside STATE_FIELDS |
StreamRowMerge::apply() |
stream.state event (5), kept in the node's own store (Core\Cluster\StreamRuntime, 7) |
Domain\Stream\StreamSource |
the stream row, this node's streams_servers row, the stream options and a recording, read before running a stream or a recording; the stream with this node's row and its runtime state (nodeRow(), workerRow(), plainRow(), createdRow(), builtServerRow(), channelRow(), movieRow()) |
SQL | the node's stream caches, built by cluster:apply from the R2 streams section once the STREAMS flow is on (ReplicaStreamCache, 7), with the runtime state from the node's own store once it is seeded (StreamRuntime, 7); stream_bundle on a miss (not built) |
Domain\Stream\NodeStreams |
the lists of this node's streams its crons and daemons select with their runtime state (cron:streams, cron:vod, cron:cleanup, the on-demand daemon) |
SQL | the stream caches and the node's own store, or the R2 section whole (ReplicaStreams) for the lists cron:cleanup prunes files by (7) |
Core\Cluster\LogSink |
client, stream, stream-error, panel-error and restream-detection records; root's system log lines (syslog()) |
one multi-row INSERT per batch (chunks of 1000); the caller's own mysql_syslog INSERT |
log.* events, redacted first (5); log.syslog (7) |
MAIN side¶
| Seam | Role |
|---|---|
Domain\Stream\StreamRowMerge |
Merges a node's runtime state into that node's row only. eventFields() keeps only runtime columns and redacts the source URL. |
Domain\Stream\StreamCacheBuilder |
The stream_<id> cache entry: columns, per-server rows, and the rule that a stream's source URLs stay out of the cache unless it is a direct source. cron:cache_engine, which only MAIN runs, builds it with this class. |
Core\Cluster\Redactor |
Strips password=, token= and username= values, the /user/pass/ segments of Xtream URLs, and user:pass@ from text before it is journaled. |
MAIN → node¶
| Seam | Wraps | Catalogue | API form (phase 4) |
|---|---|---|---|
Core\Cluster\NodeRpc |
ApiClient::systemRequest() / asyncRequest(): request/response calls to a node's /api |
NodeRpc::ACTIONS |
node.rpc{action}, answered via ack / rpc_result |
Core\Cluster\NodeActions |
root actions for RootSignalsCronJob: reboot, services, update/rollback, ports, sysctl, certbot, modules, blocklist flush, OPENSSL_EXTRA, and the credential, key and pin actions that run over the cluster API only |
NodeActions::ROOT_ACTIONS |
node.root{action} for cluster:root, with an artefact grant when the action needs a file of MAIN's |
Both seams refuse an action that is not in their catalogue, so a new call is a
deliberate change. NodeRpcActionsTest checks that every call site uses a
catalogued RPC action and that RootSignalsCronJob handles every catalogued root
action.
Rules for new code¶
- Do not write
INSERT INTO signals,UPDATE streams_servers(for runtime state), or an INSERT into a log table directly; use the seam. - Do not call
ApiClient::systemRequest()/asyncRequest()orSignalDispatcher::rootAction()outsideCore\Cluster; add the action to the catalogue and useNodeRpc/NodeActions. - Node-side code reads stream definitions through
StreamSource, not with its own queries onstreams,streams_optionsorrecordings. Once the node's replica owns the streams (STREAMS flow on), its answers come from the stream cachescluster:applybuilt, never from MAIN's database. A stream read with this node's runtime state (a joinedstreams ⨝ streams_serversrow, or a list filtered onpidorstream_status) goes throughStreamSourceorNodeStreamstoo: withStreamSource::local()(the replica owns the streams and the node's own storeStreamRuntimeis seeded) the runtime columns come from that store, which the writers keep with STREAMS on, and the seam keeps MAIN's statement for everywhere else. Do not reconnect MAIN's database in a daemon whereStreamSource::local()holds. A runtime column MAIN writes itself for this node reaches that store only where node code keeps it too (the recorder keeps the VOD MAIN attaches to the node for a finished recording). - Node-side code reads settings through
SettingsManager's getters and servers throughServerRepository. A node in mode 2, or in mode 1 with the CONFIG flow on, boots from its replica (ReplicaStage) once an apply built its caches: those come from the replica's caches, and any other query opens MAIN's database lazily, on first use. Mode 1 opens it, counted at the query's site; mode 2 refuses it. The streaming entry points take the same lazy handle there (LegacyInitializer::initStreaming). On a node in mode 1 or 2, a settings key outsidelb_settings_keys.phpis counted as a miss (SettingsAudit) and shown on Servers → Cluster Nodes. - Do not open MAIN's database before a query needs it. A connect at boot or at
the top of an entry point counts against a mode 1 node's seven-day zero even
when the request ends without a query. Where the replica may not answer
(
ReplicaBoot::hybrid(): mode 1, never where connects are refused), read MAIN's database on the lazy handle, never on a new one. To learn whether MAIN's database answers, query it: a lazy handle'sconnectedstays false until its first query (StatusCommand::mainDatabaseAnswers()). - Do not write
new DatabaseHandler(). Take the process's handle (DatabaseAware,DatabaseFactory::get()), orDatabaseFactory::connect(),connectLazy()oropen(). Every connect to MAIN's MySQL or Redis passesConnectAudit::guard(): on a node in mode 1 or 2 it is counted with its caller and shown on Servers → Cluster Nodes, and in mode 2 it throwsLbDatabaseAccessException. Do not open PDO,\Redisor mysqli connections of your own in code a load balancer runs. - A node in mode 2 (
NodeRole::refusesConnects()) never falls back to MAIN's database. A write with an agent path goes through its seam first (LogSink,LogSink::syslog()for root's system log lines,NodeStateSink), and a root action runs after its log line whatever became of the line (a line the agent did not take stays in the panel's error log). An action that still needs MAIN's database is refused up front in mode 2 (RootSignalsCronJob::updatesHere()forupdateandrollback). Work that needs MAIN's data no replica section carries yet is skipped in mode 2 behind a named seam (CleanupCronJob::streamChecks(), which holds once the replica and the node's own store answer), never run against an empty or partial answer. A stream-state write the agent did not take stays in the node's own store in mode 2 (StreamStateWriter::resend()sends it later) instead of falling back to MAIN's row; in mode 0 and 1 it falls back, and the store lapses once MAIN's row has it (StreamRuntime::lapse()). - A file a node needs from MAIN (a custom off-air video, a module's archive)
is an artefact: MAIN names it in
Domain\Cluster\ArtefactRegistryand grants it with a signed command (ArtefactGrants), and the node uses it only onceCore\Cluster\ArtefactStagechecked its size and SHA-256 against the grant, as it copies it to where it is used (root's own stage for a root action). Do not add a pull from MAIN by path, URL or password. - MAIN's work for a node goes through
SignalDispatcherandNodeActions, never asignalsrow written by hand: a node in mode 2 reads none. There MAIN sends cache jobs asnode.cachecommands inCacheJobs::job()'s form, the node runs the jobs it queues for itself at once, and a value the node read back from its own row in MAIN's database comes from the copyNodeStateSinkkeeps of what it reported (NodeStateSink::reported()). - The node's audit files (
storage/cluster/,config/cluster/audit.json) sit where xc_vm can write. A root process writes them only insideSettingsAudit::asAgentUser(), which does the work as xc_vm, never with root's own rights.
Tests that pin these rules: SignalDispatcherParityTest, StreamStateWriterTest,
StreamRowMergeTest, StreamCacheBuilderSourceTest, LogSinkTest,
NodeRpcActionsTest, ArchitectureTest, DbConnectRefusalTest,
ReplicaBootTest, ModeTwoPathsTest, ArtefactHashRefusalTest,
StreamRuntimeTest and StreamRuntimeReadersTest.