09 / Reinforcement learning

DOOMORPH

A combat policy measured across seven held-out scenarios.

Held-out suite · five episodes per scenario

Held-out / five episodes per scenario

Competence map

The values above come from the saved competency report. Gameplay footage is held for media-rights review.

Reading the suite

01

Inputs

The policy combines game pixels, stereo audio, previous action, and recurrent state.

02

Cards

Each scenario card opens its measured result from five held-out episodes.

03

Pattern

Combat and line defense pass. Movement, gathering, cover, and position prediction are weaker in this run.

DOOMORPH trains a recurrent PPO policy from game pixels, stereo audio, previous action, and episode boundaries. The visual encoder starts from random weights.

The held-out suite passes basic combat and line defense. Health gathering, navigation, cover, and position prediction miss their thresholds. Open any scenario to see the measured result.