09 / Reinforcement learning
DOOMORPH
A combat policy measured across seven held-out scenarios.
Held-out suite · five episodes per scenario
Held-out / five episodes per scenario
Competence map
The values above come from the saved competency report. Gameplay footage is held for media-rights review.
Reading the suite
Inputs
The policy combines game pixels, stereo audio, previous action, and recurrent state.
Cards
Each scenario card opens its measured result from five held-out episodes.
Pattern
Combat and line defense pass. Movement, gathering, cover, and position prediction are weaker in this run.
DOOMORPH trains a recurrent PPO policy from game pixels, stereo audio, previous action, and episode boundaries. The visual encoder starts from random weights.
The held-out suite passes basic combat and line defense. Health gathering, navigation, cover, and position prediction miss their thresholds. Open any scenario to see the measured result.