pong-ml nodejavascript.com

Both of them are playing to win. Neither of them can.

That is a reward function, not a slogan. Every ball a paddle gets back is worth one point — and a ball sent to the far end of the other paddle's side is worth two. Both sides are paid to hurt each other, and the rally gets longer anyway, because the only way to survive your own attack is to be good enough to survive theirs.

The game is the original: two paddles, one ball, a return angle you steer with where you meet the ball, and a ball that speeds up on every return. The ball is never flat — the flattest return it will give you is 22° above the horizontal, and you can send it steeper than that by meeting it with the edge of your paddle. Which is the point: aiming is what attacking is, and a paddle that only ever defends is a paddle that never wins.

3D by three.js · two 65-weight networks trained in the tab · no library for the learning · no server, no account, no keys.

Watch a pair learn to never miss

They start by doing nothing at all. Each network is wired to random noise with its output connected to nothing, so a paddle given a ball just stands there: watch the first serves of its life go past two paddles that are completely still. Then it works out that it can move at all, then which way, then how far. The paddles on this court are not a demonstration player standing in for the real one — they are the two being trained, and the numbers around the court are what they can manage at this moment. Reload the page, or press Start again, and they go back to knowing nothing.

What it is learning, while it learns it

Best score, both searches. Each side's own number, never a shared one. The dashed verticals mark where the examination got longer, which is where the number steps, and where the pair last held every serve to the end of it.
Returns a serve the pair manages, against the examination it has to pass — the dashed rule is that target and the axis does not move with the line. This is the number that has to climb.
The held-out examination — balls served with a different seed, never trained against, so this is the only chart that can say whether the pair has learned anything that carries. It is the share of a whole examination the pair holds, and the dashed rule is that whole examination. Points appear only where a reading was taken, and the vertical rules mark the generations at which the examination got longer — which is where the line steps down.
The most ground each side has demanded of the other, per return, in any examination so far. This is the attack, and it belongs to the side that made the shot — the term that makes these two learners rather than one learner with two names.
Learning

Rally now0
Returns a serve0
Longest rally0
Misses watched0
Ball speed0
Generations0
Mean gap—
Mean rush—
Two searches, two scoreboards. Neither can read the other's.
At this moment Left Right
Points won off the other 0 0
Balls it dropped 0 0
Best score 0 0
…holding unseen balls — —
Its own step size 0 0

What the machines are paid for

A policy that has never returned a ball has to be told something, or there is nothing to search along. So the reward has four parts, and all of them are printed here because a reward you cannot read is a claim you cannot check.

Which leaves an equilibrium rather than a truce. A weak pair scores badly because it misses. A pair where one side wins scores badly because the rally stops. The only way to score well is to be strong enough to survive a ball aimed at your far corner — and that is why both sides attacking produces the longest rally, not the shortest.

How they learn

By hill climbing, which is the least clever method that works and the easiest one to check.

The line you cannot cross

A word on the angle, because it is the game. The first version let a paddle put the ball back on the horizontal by meeting it dead centre — and a flat ball at a constant height is returned without moving at all, so doing nothing scored perfectly and the rally never ended. The second version floored the angle at 15°, which was still wrong in a way you can see by playing it: a paddle that meets the ball dead centre returns it AT the floor, every time, so the whole game sat at one shallow angle. The floor is now 22° and the ceiling 40°, so no return ever reads as flat and there is a real band for a paddle to aim within.

And the honest part: the floor does not change the value of the line. The ceiling sets it either way — 1,860 ÷ sin 40° is below 1,860 × 1,064 ÷ (580 × cos 22°) however flat the floor is. What the floor changes is whether the line means anything. With no floor, a flat ball is returned without the paddle moving, the ceiling limit never arises, and two careful players rally for ever at any speed at all. So the floor is not a term in the arithmetic; it is what gives the arithmetic any force. The tempting sentence — "the floor is what makes the line the real constraint" — is false here, which is why it is not on this page.

Questions this page gets asked

What are the machines actually paid for?

One thing: the distance between the paddle and the ball at the moment the ball arrives. A return scores a little; a miss costs a great deal. Nothing in the objective mentions the opponent, so no policy can earn anything by making the other side miss.

Why does a rally ever end if both sides are trying to win?

Because wanting it to continue is not what either side is paid for. The ball gets faster, and it is never flat — the flattest return it will give you is 22° above the horizontal, so every crossing makes the paddle travel and no rally can settle into a horizontal exchange that returns itself. There is a speed above which the ball crosses the court faster than a paddle can run its own length, and past that point a miss stops being a mistake and becomes arithmetic.

Does anything I do here leave my browser?

No. The game, the training and the networks all run in the page, and they are gone when the tab is closed. The site sets no cookie of its own. Google Analytics loads only after you agree to it, and refusing makes no request to Google at all.

Privacy

This policy covers pong-ml.nodejavascript.com and nothing else. It was last changed on 26 September 2026.