The Machine That Draws Never Needed a Rival
Convincing fake photos were supposed to require two programs fighting each other. One program, patiently walking an image out of static, matched them.
You have probably heard how a computer learns to invent a human face that never existed. Two programs, locked in a duel. One forges photographs; the other plays detective and tries to spot the fakes. Every failure teaches both. The forger improves, the detective sharpens, and after millions of rounds the forgeries pass. It's a good story, and it was the reigning explanation for why machine-made images got convincing. Competition was the engine. Remove it and the quality was supposed to collapse.
There was another idea on the table, and for years it looked like a dead end. Picture every possible image as a point in an enormous space — every arrangement of pixels a grid could hold. Nearly all of that space is static. Real photographs take up a fantastically small portion of it. Suppose you could build a compass that, standing at any point in that space, pointed toward arrangements that look more like a real photograph. Then you could start anywhere and simply walk to a picture. No rival, no contest, just follow the needle.
The trouble was where the photographs sit. Researchers had long assumed they lie along a trail so thin it has no thickness at all, a filament threaded through a vast empty room. On the filament, the compass works. One step off it, there is nothing to measure and the needle spins. And you always begin off the filament, out in the static, because the static is the only place there is to begin.
The move was to stop treating static as the enemy. Take real photographs and deliberately fog them with it. A lightly fogged photo sits just beside the filament. A heavily fogged one could be almost anywhere. Do this at ten strengths, from a haze that barely shows to a fog that swallows the picture, and the filament smears out into a cloud that fills the whole room. Now every point in the space has a direction attached to it. A single network is trained to read the compass at all ten fog levels at once.
Making a picture then becomes patient work. Start with pure static. Read the compass at the thickest fog setting, nudge the pixels the way it points, repeat. Turn the fog down a notch and go again. About a thousand small corrections later, the static has settled into an image.
On the standard scoreboard for this work, a collection of small images called CIFAR-10, the no-contest method scored 8.87, above ProgressiveGAN's 8.80 and above every duelling model on the list. A second measure of quality put it in the same band as the best of them. The gradual step-down was doing real work, too, not just the fog: a version given one faint fog level and no schedule produced nothing recognizable.
It is slow. A duelling model makes its picture in a single pass; this takes roughly a thousand. The fog schedule has to be tuned by hand, and the whole demonstration runs on images 32 pixels across, so how far it carries is unknown. What it does settle is that the rivalry was never the necessary ingredient. The compass was never broken. It was being asked for directions in the one place it had none, and the fix was to ruin the photographs first.