About this episode
This is the third episode of the Big Picture Medicine podcast, and I'm Mustafa. What you just heard there was a deep fake — a fake video made by Jordan Peele using an artificial intelligence technique called Generative Adversarial Networks, or GANs. Using this technique, you can feed in video of someone speaking and then generate a completely new video of them saying whatever you want.
In this episode I speak to Dr Siegfried Wagner about how these kinds of technologies can be used in medicine. We also talk about AlzEye, an unprecedented study going on at Moorfields Eye Hospital. It's looking at two million retinal images and seeing how they can be used to predict diseases such as Alzheimer's — using, of course, AI. I hope you enjoy it. It'll be particularly interesting if you're into ophthalmology and want to deep-dive into how some of these AI techniques actually work.
In this conversation
- The eye as a window into the body: from 1930s blood-pressure survival stats to a 2018 Google Brain model that predicts your age — and your biological sex — from a single retinal photo with 98% accuracy.
- What AlzEye actually is: a record-level linkage of ten years of Moorfields retinal images (2008–2018) to nationwide hospital data, so you can find people who were well when scanned but later developed dementia, a stroke or a heart attack.
- Why screen for Alzheimer's when there's no cure? Wagner's case: three-quarters of people see an optometrist, most want to know their risk, and a five-to-ten-second OCT could become a cheap, non-invasive clinical-trial endpoint.
- Cracking open the "black box": occlusion testing and saliency maps to see what a model is really looking at — and the confounders (poorer-quality scans, more cataract) that can quietly fool you.
- The legitimate side of deepfakes: GANs for cross-modality transfer — turning a CT into an MRI, or inferring 3D structure from 2D images — and the hard problem of knowing how faithful the synthetic image really is.
Transcript AI-generated
How much can you tell about someone from looking at their eye?
“In 2018 there was a big paper from Google Brain showing that from a single retinal photograph you could predict a person's age — and their biological sex, with 98% accuracy. We didn't know you could tell that from the back of the eye.”
Siegfried
Classically, we know you can tell about poorly controlled disease states — the typical one being high blood pressure. Most people know that blood pressure can affect the back of the eye, and it does so in a diverse set of ways. You can even use those features to stratify people by their risk of death: back in the 1930s, people could look at the blood vessels at the back of the eye and find a really strong association with survival rates. That's blood pressure. Diabetes is obviously another classic thing that affects the eye — that's associated with diabetic control, but also with other features like blood pressure and smoking.
And then what's emerging more — because what I focus on are the real globally leading causes of morbidity and mortality — is dementia, neurodegenerative disease. That's a more recent understanding. When I say recent, I mean just the last couple of decades, and really only the last decade. There's even a little bit of evidence that you can use retinal features for psychosis, and in particular for the development of schizophrenia. So there's lots you can tell.
That's all classical, traditional epidemiological analysis: I know these people have this disease, I look at the association. But now that we have large amounts of high-dimensional data — things like imaging of the back of the eye — we can start using more modern techniques. The thing everyone's reading about is these new artificial intelligence methods, things like deep learning. And we've found that using deep learning, we can derive even more about people just from the back of their eye, from simple retinal photographs.
In 2018 there was a big paper from Google Brain which showed that from a single retinal photograph you could predict, with impressive accuracy, a person's age — and their biological sex, with 98% accuracy. What's really interesting is that we've known for some time that as we get older, certain things change at the back of the eye: the arterioles get more tortuous, the calibre changes. So it's not surprising that deep learning can tell us someone's age — perhaps surprising it can do it so accurately. But we didn't know you could tell someone's biological sex from looking at the back of their eye. The point is that these methods have the potential to give us things we already know, but in a more accurate way, and also to spawn this area of discovery science: does it identify new biomarkers we hadn't considered before?
So the study you're working on is called AlzEye. What exactly is it, and what are you looking for?
To give you some context: a lot of the work we've done, and a lot of the papers I've mentioned, rely on these large, expensive prospective epidemiological studies — things like UK Biobank, which cost tens of millions of pounds and are incredible for the amount of data you have on each individual. The problem is that these are often healthy cohorts. They rely on patient recruitment, the populations can be quite homogeneous — Biobank, for example, is over 92% Caucasian, and they're healthy people.
So one other option is to try to leverage real-world data, and that's what AlzEye tries to pioneer. AlzEye is a record-level data linkage set. It links images taken at Moorfields Eye Hospital between 2008 and 2018 — so ten years of retinal images, that's retinal photographs and something called OCT, optical coherence tomography — with national data on hospital episode statistics: admissions data, outpatient clinics, A&E attendances, on a nationwide level. Any time you go to a hospital and you're admitted for day-case surgery or an inpatient stay, that's coded within hospital episode statistics. The coding has improved in quality particularly since 2004, and through a long series of approvals you can retrieve that data. To the best of our knowledge, it hasn't really been linked at high volume with ophthalmic imaging before. That's what AlzEye sought to do.
Pearse Keane is my supervisor and mentor, and we spent the last two years getting the necessary approvals in place to duplicate the images at Moorfields and link them with hospital episode statistics. What that means is you have a dataset of patients who may have had a retinal photograph in, say, 2012 — and they were fine, they were well — but by 2018 they'd developed dementia or Alzheimer's disease, or had a heart attack or a stroke. So we've got imaging that predates that. The goal is: can we establish that there's some predictive value in these images that will help us identify the patients who are most at risk?
So you're taking these retinal images, looking at patient outcomes later on, and trying to draw a connection.
That's right, yeah.
And why is it specifically Alzheimer's that you've looked into?
We've known for some time that people who have Alzheimer's disease have thinner nerves at the back of their eyes — something called the retinal nerve fibre layer. In cross-sectional data, if you measure the back of the eye in someone with Alzheimer's against an age-matched control, there's a significant difference. But what's really emerged in the last 15 to 18 months — very convincingly — is that these findings are not just indicative of prevalent dementia; they also predict cognitive decline and dementia. That's come from UK Biobank and from the Rotterdam study, and interestingly they were published in the same issue of JAMA Neurology, I think October 2018. UK Biobank showed that patients who perform more poorly on the mini-mental state examination have a thinner retinal nerve fibre layer — and, in about 1,250 patients, that those with a thinner layer are also more likely to do worse on that cognitive questionnaire a few years later. The Rotterdam study shows something similar, but now looking at actual labels of dementia — so it's not just cognitive decline, which is one thing but doesn't necessarily mean someone has a neurodegenerative disease. They show it in cases of dementia.
How do you see your results being used clinically — for screening, or as a diagnostic tool?
“It's a slightly controversial area, because we don't have effective treatments for dementia at the moment. A lot of people would say, well, what's the point of picking up a condition with a screening test if you don't have a treatment for it?”
Siegfried
That's really important. It would be great if it were a diagnostic tool, but I think that's very unlikely. What's more likely is a screening situation. We're talking here in the UK, where we have a particular infrastructure for eye health, mainly through community optometry. If you've listened to the news — I probably shouldn't say the name — but the largest franchise in the UK for community optometry now has an OCT device in every single branch. So high-resolution imaging of the back of the eye is becoming ubiquitous, and there's a role for screening there.
Back in 2012 you may have heard of the over-40 health check for cardiovascular disease. At the time, I think around a third or less of people actually attended. Uptake has improved since, but it's still a smallish number — whereas more than half, even around three-quarters, of the population will attend an optometrist, because vision is a very important sense. There's some work from a collaborator here, David Crabb at City University, which showed that patients rate vision above the other senses very highly — I think they even showed a patient would be willing to give up five or eight years of life to maintain their vision. That's how important it is. So there's an ability to target a group that might not otherwise go to their GP just for some risk stratification for cardiovascular disease or dementia.
It's a slightly controversial area, because we don't have effective treatments for dementia at the moment. A lot of people would say, well, what's the point of picking up a condition with a screening test if you don't have a treatment for it? This goes through the Wilson criteria for screening. I have a lot of responses to that. But screening is one objective; another is whether this could be a useful biomarker for clinical trials — a clinical-trial endpoint. One of the challenges with some of the drug trials for dementia is that they often don't pick up patients early enough, and the endpoints are highly invasive — the level of amyloid on CSF via lumbar puncture, or an MRI/PET amyloid protocol, which is expensive, labour-intensive and requires expertise to analyse. So could something like a non-invasive OCT scan, which takes five to ten seconds and is relatively cheap, provide an alternative? That's another question.
If there's no benefit to early intervention and no treatment, why screen for Alzheimer's at all?
The first thing is that it's screening, not diagnosis. Someone could then go on for further testing — and we might find, well, you have a high risk, or you have a lot of vascular disease, and there are many things we can do to optimise that. So it's not just about a risk of dementia; it's about the presence of risk factors. Another point: when you survey people who have a family history of dementia, they want to know if they're at risk. There's a really huge survey out of the US, from one of the Alzheimer's research organisations there, and it's something like over 80% of people who would want to know even if they were at only slightly higher risk. There is some weak-to-mild evidence for pharmacological treatments in the early stages of disease, and then there are all the other measures — it can encourage smoking cessation, exercise, improving your diet. But I acknowledge it is a controversial area; we don't yet have a treatment. The other hope, of course, given how much money is going into dementia research, is that we do have something in the future.
On AI, there's been a lot of talk of black boxes — the argument being that you put these retinal images into some algorithm and it comes back with a yes or no, "this person is going to develop Alzheimer's," and you wouldn't know how it reached that decision. Does that apply to what you're doing, and are there ways of finding out what the algorithm is looking at — how you know it's looking at the retinal nerve fibre layer and not something else?
This is very much my opinion, but in our approach the first step is traditional statistical modelling. You segment something like an OCT — measure someone's nerve fibre layer, for example — and see how that associates with the development of dementia, so you get an idea of how strong that signal is. Then, in later steps, you employ something like deep learning, which is particularly powerful for medical imaging data, and see what incremental value that model gives over the traditional association. And, as you say, you want to see what it's picking up. Is it picking up noise? For example, people with cognitive impairment might give poorer-quality scans because of more movement artefact, or they might be more likely to consult eye services and so have more cataract. So we have to be very cognisant of those potential confounders.
There's a lot of work going into solving the black-box problem. Nothing's really solved it, but there are techniques that look at the interpretability of these models. One used especially in the ophthalmic world is occlusion testing: you occlude certain parts of the image and see how the accuracy of your deep-learning model changes. There's a nice paper in Nature Biomedical Engineering, again by the Google Brain group, on the detection of anaemia from a fundus photo. What they find is that occluding a lot of the fundus photo doesn't particularly affect model accuracy — but when you occlude the optic nerve, accuracy drops off very quickly. That tells you the optic nerve is obviously crucial to the decision-making. You can do something similar on OCT scans or any other medical imaging.
So occlusion means you cover up parts of the image — and if your model stops working, you know it was probably using that part?
Exactly. In that study they used many different occlusion techniques. One is to occlude the peripheral retina — you black it out so all you can see is the macula, the optic nerve and the major blood vessels around the posterior pole — and you see the accuracy is still very good, so the peripheral retina is probably not that influential in the model's decision. Then you can cover the vessels, then the macula, then the optic nerve. There are other techniques too — saliency maps, which show you the pixels that contribute most. This whole field is called attention: where is the model looking, what's affecting its decision-making? And there have been other advances in the field, like GANs — generative adversarial networks.
What impact could they have in ophthalmology?
“Most of the stories people read in the news about AI and deep learning — "as good as doctors," "exceeds the performance of doctors" — are discrimination or classification models.”
Siegfried
Just to quickly explain what GANs are. Most of the stories people read in the news about AI and deep learning — "as good as doctors," "exceeds the performance of doctors" — are discrimination or classification models. You put an image into the model and it comes out with a decision: I classify this into this category, this is the diagnosis. But there's another type called a generative model — models that produce or construct data. There's a lot of potential here. The most obvious application, the one that gets a lot of hype, is synthetic images: these can generate images that have never been seen before, that don't exist in the real world. You may have seen stories where a GAN is fed loads of images of celebrities and generates new images of people who look feasible — some look slightly abnormal, and you can tell they're synthetic from certain pixels, but generally they produce faces that don't actually exist.
Right — and there are videos of politicians.
That's right. Even today there was some of that in the news — these deep fakes and how dangerous their potential is. That's one application, but there are many. You can use GANs for something called cross-modality transfer: people have input images from one imaging modality and output a different one. Can we extract data the human eye can't see? With a GAN you could take, say, an anterior and a lateral x-ray view and form a CT from them. There's work looking at going from a CT to an MRI, and efforts to gauge three-dimensional information from two-dimensional imaging. It's challenging — the difficult thing is assessing how faithful the process is. How much do the generated images actually resemble the real ones, and how can clinicians tell the difference? There are many different ways to assess that.
I don't want to embarrass you, but you've obviously had quite a successful academic career. I ask everyone this: along your journey, have there been habits, ways of thinking, resources or books you'd recommend to someone who wants to follow a similar path?
That's very flattering — I'm not sure it's that glowing. One thing I've always benefited from, from medical school to now, is really inspirational and supportive role models. It's important to identify role models at every stage of your career and to spend as much time as you can with them, without frustrating them — because they'll provide enormously helpful feedback, and you can learn a lot from how they've proceeded in their lives and their academic careers. I've benefited from inspirational role models at medical school, through my early training, and even now in my PhD. And they don't need to be in the area you're interested in — a lot of these skills transcend specialties.
I hope you liked that episode — make sure you subscribe if you did. If you have any thoughts or feedback, the best way to get in touch is on Twitter, @MustafaSultan. Links to everything mentioned can be found in the show notes. Thank you.