REINFORCEMENT LEARNING, WITH A JUDGE
Let JEV judge.
Let the agent learn.
A tiny agent. A world to figure out. Every move becomes a lesson.
Watch a reward signal turn trial and error into a winning run.
OBSERVE→ACT→JUDGE→LEARN
One decision at a time. ↲
01 / THE PLAYGROUND
TRAINED POLICYKey Quest 9 × 7
STEP 00FIND THE KEY
TRAINING TIME MACHINE
Agent Key Exit LavaCollect the key. Avoid lava. Reach the exit.
04 / THE LEARNING CURVE
Practice leaves a trace.
Last 20 episodes · success
EXPLORATION ε 0.04350 EPISODES
05 / THE RESULTSame world.
Same world.
Better decisions.
BEFORE—
→AFTER—
Success on 60 evaluation episodes, same seeds and map. Evaluation makes no judge calls.
The judge supplies the reward. The agent learns the actions. Tabular Q-learning updates from the judge’s outcome probabilities. JEV receives structured state, including computed route distances; the agent sees only position and key possession. This is a small, inspectable experiment—not a pixel-based or general-game benchmark.
About JEV Score ↗