MP-PPO leads the balanced case.
Strong voltage control, 98.9% critical supply and healthy terminal SOC.
Best portfolio resultMaster Project 02 · Research case study
Phase 6 · G8 passFour cooperative policies stabilise an islanded 5.5 V DC bus, protect its battery and prioritise a critical load — tested against droop and MPC across a frozen Phase 6 release.
No universal winner. Learned control leads the balanced slow case; domain randomisation strengthens default; classical droop remains the better fast reference.
Vbus 5.5 V
01 / From hardware to control
MP1 builds the physical energy node. MP2 turns that setup into a physics-based digital twin and asks a sharper question: where can learned source controllers beat droop and MPC — and where do the classical controllers still deserve respect?
Bus voltage is not clamped by an ideal supply. It moves with the net source current through the 7.5 F buffer, so voltage becomes the visible consequence of every control decision — the DC analogue of frequency in a larger AC grid.
Strong voltage control, 98.9% critical supply and healthy terminal SOC.
Best portfolio resultLower RMSE and saturation with better supply and reward across held-out seeds.
Cleanest DR gainThe short coupled scenario resists a simple “faster RL is better” conclusion.
Honest boundary02 / Plant model
The simulation integrates bus physics at one-second substeps, while the controller acts at 30 s, 5 s or 1 s intervals. Slow and default use 48-hour episodes; the fast setting has a shorter horizon and scenario coupling, so cross-configuration differences are reported as configuration effects — not as a pure action-interval causal claim.
Passive source, temperature-adjusted and curtailed by its policy. Phase 6 is bypass-only.
Cubic turbine curve from cut-in through rated speed to cut-out.
Dispatchable source with ramp, minimum runtime and cooldown constraints.
0.5 Ah bidirectional storage with SOC limits and 95% charging efficiency.
03 / Learning setup
The headline controller is best described as Multi-Policy PPO: four PPO policies receive the same 21-dimensional global observation and optimise one cooperative reward. It is not the original centralised-critic MAPPO architecture, so the project names the implementation precisely.
Behaviour cloning from droop provides the warm-start. The Phase 6 matrix then compares it with local-observation IPPO, single-agent PPO, constrained Lagrangian PPO, domain-randomised training and two inference-time shield settings. Classical references include droop, heuristic MPC with three forecast-noise levels and convex MPC.
Why the gate structure mattered. Single-agent PPO looked attractive during training, then failed to generalise: in the default held-out evaluation its critical supply collapsed. Phase 6 reports the held-out test result, not the flattering training curve.
04 / Experiment scale
The main Phase 6 campaign trained five controller families at three action intervals and five independent seeds: 75 complete jobs, 200 iterations each. Their measured wall-clock spans sum to 156.8 hours; the matrix itself ran from 14 to 25 June 2026.
The final T7/G8 bundle then adds 120/120 regular evaluations, 108/108 stress evaluations and 24/24 behaviour traces. A headline Multi-Policy PPO run took a median 2.41 hours; the faster single-agent PPO was not the better held-out controller.
five algorithms, three action intervals and five seeds
the health-checked headline training matrix
iterations in the reward-weight sensitivity sweep
regular evaluation, stress scenarios and behaviour traces all released
05 / Phase 6 results
In the completed slow configuration, raw Multi-Policy PPO reduces mean voltage RMSE from 0.554 to 0.291, raises critical-load supply from 65.6% to 98.9%, and ends with mean battery SOC 0.786 instead of 0.052. The domain-randomised arm improves RMSE further to 0.286, but slightly lowers critical supply, final SOC and reward — a real trade-off, not free dominance.
In default, domain randomisation improves RMSE, saturation, critical supply and reward over the raw learned arm. In fast, droop is stronger than the DR arm on the main voltage, saturation, supply and reward metrics. The final page therefore presents Phase 6 as a configuration-dependent result instead of a universal “RL wins” claim.
Forecast robustness
With a perfect forecast, MPC is competitive and in some metrics better. At the Phase 6 15% noise setting, however, its RMSE rises to 0.548 and critical supply falls to 68.4%. The learned controller has no forecast input, so this specific error channel does not exist for it.
06 / Original Phase 6 evidence
The summary graphics above explain the result quickly. This browser keeps the original Phase 6 plots available underneath: three Holm-corrected effect-size matrices and five held-out behaviour traces. Every trace uses the same unseen EVAL seed 100, so the controller families face one directly comparable scenario.
8 / 11 release figures shown The three original MPC-noise exports are represented by the cleaner forecast-robustness graphic above; the underlying conclusion and caveat remain unchanged.
07 / What Phase 6 established
08 / Technology
Ergebnisse / Phase 6ReleaseT7 complete · G8 passRelease commit15befd5Bundle frozen17 Jul 2026