The Radiology AI That Can Finally Point to What It Sees
AI can already write your scan report. The problem is nobody knows where it was looking — and that may matter more than whether it's right.
A radiologist reads a scan. Then she circles. Faced with a CT image, a cross-section of the body rendered in shades of gray, she writes down what she finds, but the writing is only half the work. The other half is the pointing. Here, she says. Not approximately, not somewhere in this neighborhood. Here. The circle is a promise that the words can be checked against the picture.
For several years now, artificial intelligence has been learning the first half of that job. Feed a modern system enough scans and reports and it will write you a fluent, plausible summary of what an image contains. None of these systems could draw the circle.
That sounds like a refinement to bolt on later. It isn't. When a machine writes "the lower right lung shows increased density" and a doctor has no way to see where the machine was looking, the sentence stops being a finding and becomes a guess in a confident voice. To trust it, the doctor has to go back and re-read the entire scan from scratch, at which point the machine has saved no one anything. In medicine, a claim you cannot locate is a claim you cannot use.
So why had no one simply built the pointing in? Because teaching a machine to point was understood to demand an enormous, tedious favor from people. Someone, a radiologist or a trained technician, would have to sit with thousands of images and hand-draw a box around every organ, one at a time. For CT and MRI scans, which are three-dimensional and take far longer to mark up than ordinary photographs, those hand-drawn labels barely exist. The field had mostly accepted this as a wall: the pointing would have to wait for data that no one had made.
A team at a university hospital noticed the labels were already there. Radiologists had been describing where things were for a decade. "Lesion in the left lobe of the liver." "Enlarged right adrenal gland." These are spatial claims, written in ordinary clinical language. The researchers built a system to pull those location words out automatically and match them to a ready-made map of normal human anatomy, a digital atlas of 121 body structures. No one drew a single box. Out of that came 1.2 million scan-and-text pairs, 236,000 of them carrying genuine spatial grounding, assembled essentially for free.
The model trained on this, which they call RadGrounder, writes reports and answers questions about scans as well as purpose-built medical AI. Making it point at what it described did not make it any worse at the describing. The accountability came at no cost to the thing the system was already good at.
There is a catch. The atlas it learned from maps only normal anatomy: liver, lungs, kidneys, spine. It does not map disease. RadGrounder can point to your kidney; it cannot yet circle the tumor growing inside it. And the tumor is the whole reason anyone ordered the scan. What the work proves is a principle, not a finished tool: that a medical AI could be built never to make a claim it can't locate. Whether that promise reaches from healthy anatomy to actual disease is the question still open.
Can future models be trained to consistently and precisely ground the exact pathological findings—such as a lesion or consolidation—rather than just the broader anatomical structures that the current system localizes?