Elastic Pool — occupancy, admission, scale-up, compaction
Version: 1.1.0 Status: SOURCE OF TRUTH (v1.1 ratified 2026-07-11) Last Updated: 2026-07-11
Overview
The commander places logged users on pool workers and grows each group’s pool on measured pressure, not head counts. There are no per-user caps: an idle user costs ~1 MB and ~0 cpu (measured on a real GenroPy worker), so counting heads answers the wrong question. The architecture is four stages:
1. MEASURE 2. TRANSPORT 3. INTERPRET 4. DECIDE
(worker, (/occupancy (commander: (commander:
every 5s) on the channel) window → number) policies)
The worker is a sensor: it reads raw quantities and pushes them. It computes no percentage and passes no judgement.
The commander is archive and judge: per worker it keeps a window of the last readings; the window absorbs anomalies, and time rules read directly from it — anti-flapping IS the window, not an extra mechanism.
The sensor report (fixed shape)
Every occupancy_interval (5s) the worker pushes occupancy_report() on
/occupancy:
Field |
Source |
Role |
|---|---|---|
|
register counts |
drift sentinels only — decisions use the commander’s surface, which is fresher |
|
users with a non-system RPC within the window |
activity (the touch happens at route resolution; |
|
dispatch executor metrics |
saturation signal |
|
|
process cpu fraction (None on first tick) |
|
|
memory, bytes |
|
pending datachanges, outbox size |
drain sentinels |
|
|
machine-level, diagnostic only (not attributable to one worker) |
The commander archives each report in worker_metrics — a
deque(maxlen=METRICS_WINDOW=60) of rows {ts, report, forward}, where
forward snapshots the cumulative forward counters (requests, errors,
seconds) the commander clocks around every relay.
The evaluator
OccupancyEvaluator (owned by the commander) turns a worker’s window into an
occupancy in [0, 1]:
per-row components:
memory= rss / configured limit (present only when the reading has an rss ANDmemory_limit_mbis set),cpu= min(cpu, 1.0) (the measured wall is one core — the GIL),executor= busy/total;each component is FIRST averaged over the last
SMOOTHING_ROWS = 6rows (~30s — transient spikes dominate a naive read), THEN the occupancy is the MAX of the averaged components: the bottleneck answers “can I take another user?”, an average across incommensurable resources does not;a worker with no rows (just born) reads 0.0 — it admits.
It also serves the monitor: history_of (per-row occupancy across the whole
window, the histogram), rates_of (rps and latency from forward-counter
deltas), components_of.
Measured baseline (real GenroPy worker, 2026-07-10): RSS ~34 MB at boot, ~100 MB after the first page (lazy imports + ORM), +~1 MB per connection, under pressure cpu ≈ one core (GIL) at ~35 rps with the executor saturated — cpu and executor are the true saturation signals; scale-out by process is the correct answer to the GIL.
Policies
Placement (reception-first): every user is born on the group’s reception (its FIRST routable worker) — guests live there, and a login happens where the guest already is. On login the reception KEEPS the user while its own occupancy is under
reception_threshold(default 0.5): keeping is the null action. Over the threshold it PASSES: the user is assigned among the OTHER workers (rule 2). The sole worker of a group always keeps (a group of one cannot refuse a login). Users the reception kept stay kept — re-placing them later is compaction’s business, not placement’s. (Replaces the v1.0 binary rule that excluded the reception whenever the group had more than one worker:reception_threshold=0reproduces it,=admission_thresholdmakes the reception a plain worker.)Admission (the pass): among the routable workers other than the reception, pick the one with the MINIMUM occupancy under
admission_threshold(default 0.8). None under threshold → place on the last one with a warning, and let check_capacity grow the group.Scale-up:
check_capacityspawns a worker when the group cannot place well: no non-reception worker underadmission_threshold— or, in a group of one, the reception over its keep-threshold (the moment it first needs to pass and has nobody to pass to). Guards: a tracked child not yet routable is a spawn in flight (wait for it);max_workerscaps the group (None = unbounded).Scale-down = compaction: governed by the capacity ledger (next section). Drain the least-occupied non-reception worker onto the survivors with commanded moves (see “Moving a live user”), then retire it.
min_workersis the floor (default 1: the reception). The reception is never compacted.
Placement and admission evaluate on demand, on the request path (a login
arriving); the capacity ledger is checked when an /occupancy report arrives
— the periodic signal that already exists. No commander timer either way.
Compaction — the capacity ledger
Each worker contributes its threshold as useful capacity: the reception
reception_threshold (0.5), every other worker admission_threshold (0.8).
For a group of N workers:
C = 0.5 + (N−1) × 0.8 total useful capacity
O = Σ occupancy(worker) measured, from the evaluator
H = C − O headroom
compact when H > 1.5 × 0.8 margin: one and a half workers' worth
Example (thresholds read as points): 4 workers → C = 290. At O = 230, H = 60: nothing to do. At O = 150, H = 140 > 120: compact — drain the least-occupied outer worker, retire it → C = 210, H = 60: stable.
Two properties fall out of the 1.5 margin:
Hysteresis built in: retiring a worker costs 0.8 of headroom, so post-compaction H stays over 0.4 — never near the scale-up trigger (H ≈ 0). The two policies cannot chase each other.
Moves always fit: free space on the survivors after the retirement is
H − 0.8 + o(drained), i.e. at least0.4 + o(drained): the drained worker’s users always fit, with half a worker to spare.
A worker already at zero users is trivially the minimum and its drain is an
empty move-set — the passive case needs no rule of its own. The check repeats
while H stays over the margin. reception_threshold, the margin and
min_workers are commander parameters (constructor today, per-group config
when that lands — the destiny already noted for SMOOTHING_ROWS).
Moving a live user
The move of user U from worker A to worker B is a quiesce–snapshot–switch,
orchestrated by the commander over the pool channel (send_to):
1. FLAG U marked "moving" on the surface: new requests for U are HELD
on the commander (await with a hard timeout), not forwarded.
2. QUIESCE wait until U has no live request. The forward path keeps a dict
of OPEN requests (id → {user, worker, ts}); quiesce = no entry
for U. A dict, not a counter: a drifted counter never heals,
a stale entry is visible and swept by its timestamp.
3. SNAPSHOT ask A for U's package: the user's register slice, its connection
and page entries, plus U's undelivered mailbox deposits — one
pickled blob (safe: compaction stays within a group, one runtime).
4. INSTALL send the package to B; B answers ok once installed.
5. SWITCH the surface re-points U → B, the flag drops, held requests
release toward B.
Failure semantics: B not answering within the timeout → abort, the flag drops, U stays on A (the next ledger check retries). A dying mid-move is a normal worker death: U’s state dies with the process and the surface sweeps. Guests never move — they live on the reception, which is never compacted.
Per-user consumption and the monitor
The commander attributes every forward it clocks to the user resolved from
the sticky cookie: user_consumption[user] = {requests, seconds},
cumulative, forgotten when the user leaves the surface. Exposed by
/_commander/population — an async route that fans out to every routable
worker’s /_monitor/state with a DEDICATED httpx client (monitor traffic
never mixes with the forward pool: httpx connections are loop-affine) and
fuses consumption into the user rows. A worker that does not answer degrades
to an error row, never an exception.
monitor_state() adds metrics (occupancy, components, history, rates,
forward counters, all from the evaluator); the server monitor page renders
per-app tabs — state, metrics (occupancy bars + SVG histogram), population.
Group = runtime version
Groups exist to serve different app versions (green/blue/canary) on
different runtimes; elasticity always scales WITHIN a group. Vehicles: a
venv per group (python= on the group config node) is supported today;
container spawn (podman/docker) is a door kept open — the channel already
does TCP and the announce carries a reachable host as data.
Design history
The ratified contract (requirements, measured numbers, decision log) is
temp/worker_occupancy_elastic_pool.md (v1.0, 2026-07-10; the temp draft was
removed once distilled here).
v1.1 (2026-07-11, decided in session): reception-first placement with a keep threshold (replacing the binary reception exclusion), scale-down as compaction driven by the capacity ledger with the 1.5-worker margin, and the live-user move protocol (flag → quiesce on the open-requests dict → pickled package → install → switch).