Researchers developed a new method that quickly and accurately compares a patient’s preoperative 3D medical scan with X-rays taken after operation. Clinicians may find it easier to accurately pilot less invasive surgical instruments with this technique, which could result in quicker and safer treatments.
Real-time X-rays are used by clinicians to guide instruments such as catheters and endoscopes through small incisions during numerous minimally invasive procedures. However, because X-rays are flat images, it might be difficult to pinpoint the precise location and orientation of surgical instruments within the patient’s body, which raises the possibility of difficulties.
Clinicians may manually match X-rays with preoperative 3D medical images, such as CT scans or MRIs, to aid in the localization of surgical instruments. In practice, artificial intelligence techniques intended to expedite this process are impractical since they are unable to align photos consistently for every patient.
The AI model in this new system, created by researchers and medical professionals from MIT and partner organizations, can adjust to each patient in as little as five minutes. In a couple of seconds, the model automatically matches one patient’s X-rays to 3D scans with sub-millimeter accuracy.
Known as XVR (X-ray volume registration), it performed an order of magnitude better than current AI techniques for a variety of patients, body parts, and medical procedures.
Most Americans reside more than an hour distant from a facility capable of performing noninvasive operations, such as emergency stroke interventions. In stroke time, an hour is a huge amount of time. According to Vivek Gopalakrishnan, a postdoc in the MIT Computer Science and Artificial Intelligence Laboratory (CSAIL), a recent graduate of the Harvard-MIT Program in Health Sciences and Technology, and the lead author of a paper on xvr that was published today in Nature, “making these procedures easier by combining 2D and 3D information enables these types of highly specialized life-saving procedures to be more accessible to much broader parts of the population.”
Neel Dey, a former postdoc in the Medical Vision Group who is currently an investigator at Harvard Medical School and Massachusetts General Hospital, and his advisor Polina Golland, the Sunlin and Priscilla Chou Professor of Electrical Engineering and Computer Science (EECS), a principal investigator in CSAIL, the head of the Medical Vision Group, and co-senior author of the paper, join him on the paper. David-Dimitris Chlorogiannis, a researcher and clinician at Harvard Medical School; Andrew Abumoussa, a neurosurgeon at St. Luke’s Marion Bloch Neuroscience Institute; Anna M. Larson, a pediatrician at Shriners Children’s Hospital; Nazim Haouchine, an assistant professor of radiology at Harvard and Brigham and Women’s Hospital; Darren B. Orbach, a doctor and scientist at Boston Children’s Hospital; and Sarah Frisken, an associate professor of radiology at Harvard.
Increasing The Informational Value of X-rays
Clinicians introduce instruments via a small incision during various minimally invasive surgical procedures, such as angioplasty to clear clogged arteries, and use a high-speed transportable X-ray scanner to create images that enable them to see the procedure from any angle.
However, physicians must align real-time X-rays with the patient’s preoperative MRI or CT scan in order to guide surgical instruments without unintentionally harming adjacent tissue. They are able to ascertain the tool’s location in respect to anatomical structures thanks to a procedure known as registering.
“To observe blurry, 2D images and comprehend how everything is oriented, a clinician must undergo decades of training. Gopalakrishnan says, “We want to make these 2D X-rays more informative so it becomes safer and easier to do these life-saving procedures.”
The doctor must estimate the location of a surgical tool by entering numbers into a computer or clicking anatomical landmarks on a screen using laborious and inefficient manual registration techniques.
Researchers are creating AI models that can forecast 2D/3D registration in order to expedite the procedure. However, because human anatomy varies so much, a model that suits one patient may not suit another.
According to Gopalakrishnan, it is challenging to train a deep learning model that is resilient enough to adjust to a large number of patients due to a shortage of high-quality annotated medical image data.
Instead of attempting to create a machine-learning model that could be used for every patient, the researchers created a model that was incredibly well-suited to the particular patient.
Gopalakrishnan continues, “We customize this one particular model for this one particular patient, and it doesn’t matter if it works on other people because there will be different models for those people.”
Machine Learning Tailored To Each Patient
Xvr creates thousands of synthetic X-rays from various angles using a patient’s preoperative 3D scan, such as an MRI or CT, generating roughly 1,000 images every second. To guarantee that these artificial images are realistic, it simulates the X-ray process using physics.
This physics simulation is fully based on the patient’s CT scan or MRI, as opposed to some forms of generative AI that create data from nothing. There is no space for hallucinations because xvr generates patient-specific data in a strictly physics-based manner, according to Gopalakrishnan.
An AI model that can precisely align this patient’s 2D X-rays with their 3D picture scan in a couple of seconds is trained using the xvr framework with these simulated data.
However, even though such a registration model is quite accurate, it would be impossible to use in an emergency because it would require 12 hours to train from scratch for every patient. The researchers utilized xvr to pretrain a more adaptable AI system, known as a foundation model, that can swiftly adapt to each new patient in order to speed up the process.
Over 2,000 individuals with a variety of ages, image modalities, and geographical locations provided whole-body 3D medical scans. These varied data were used by Xvr to create artificial X-rays and train a foundation model for 2D/3D registration.
This pretrained model performs registration with the same accuracy as if it had been trained from start, and it can adjust to a new patient in roughly five minutes.
According to Gopalakrishnan, “you can now get patient-specific accuracy but also in a very rapid time frame.”
Using data from five hospitals covering dozens of bones and organ systems in adult and pediatric patients, the team tested the model on the largest collection of actual 2D/3D registrations.
While running quickly enough for emergency surgery, Xvr greatly exceeded previous AI-based techniques in terms of accuracy and resilience. Additionally, the approach could be applied to enhance robotic surgery technology.
In the future, the researchers intend to concentrate on accelerating XVR for real-time deployment, carrying out additional research to confirm its dependability in different settings, and expanding the system to manage more complicated scenarios, such as moving body parts.
We have been meticulously creating and testing this algorithm over the last two years. In order to transform this study into practical instruments for navigation or deployment, we are currently working closely with surgical robotics businesses and clinical organizations, according to Gopalakrishnan.
The National Institutes of Health (NIH), the MIT CSAIL-Wistron Program, the MIT-IBM Computing Research Lab, the MIT Jameel Clinic, the MIT Health and Life Sciences Collaborative, and the Chou Family Transformative Research Fund provided some funding for this work.

