Mission 66 // October 4, 2021

Bias in Medical AI

Why the fix for biased medical AI isn't less AI — it's training algorithms to learn from patients, not from doctors.

ZO Ziad ObermeyerAssociate Professor of Health Policy and Management, UC Berkeley School of Public Health
Bias in Medical AI
0:00 // 47 min

About this episode

Ziad Obermeyer is Associate Professor at the Berkeley School of Public Health, where he researches and teaches on the intersection of machine learning and healthcare. Some of his most interesting research focuses on algorithmic bias, and how we can better build AI systems that avoid perpetuating and falling into some of these traps. We talk about his fascinating story, how he created an AI algorithm that actually reduces bias and supersedes human performance, and some of the things he's learned along the way. I hope you enjoy.

In this conversation

  • The knee-pain mystery: Black patients report far more pain than white patients with identical-looking knee x-rays. Obermeyer's team trained an algorithm to listen to the patient instead of the radiologist — and cut the "unexplained" racial pain gap by nearly half.
  • The bias hiding in plain sight: one family of algorithms decides care for 150–200 million Americans a year, and it was massively racially biased — because it used healthcare cost as a proxy for health need. The fix was retraining it on need, not spend.
  • Why more AI isn't the enemy of fairness: the counterintuitive lesson that the way to remove bias is to train algorithms on ground-truth patient outcomes rather than on human judgment — flipping AI from a tool that entrenches inequality to one that fights it.
  • A masterclass in problem-first thinking: why "I have a cool convnet, what can I point it at?" produces beautiful, useless papers — and why the best medical AI starts from a real problem, like COVID testing at the Greek border or heart-attack triage.
  • The 40 vs 40,000 variables insight: doctors are shockingly close to an optimal model — for the handful of variables they can hold in their heads. Where algorithms win is holding 40,000, cutting one-year mortality ~30% for high-risk patients doctors currently under-test.

Transcript AI-generated

Musty

Would you mind telling me a little about your story, and how you got to where you are today?

Ziad

Sure. It was mostly a random walk — but the nice thing about looking back over a random walk is that there are lots of stories you can tell ex post.

I started off at university studying history, actually. It was history of medicine, so it was somewhat relevant, but I'd always been interested in how science gets made — the sometimes arbitrary or social processes that go into making science. So I was approaching medicine at that point from a social science point of view. I did a master's after university in the history and philosophy of science.

But at some point — I wouldn't say I got frustrated — I just wanted to do something rather than write about it. The first thing I tried was working at a management consultancy, for nine months. Even though I learned a lot and made a lot of friends I'm still in touch with today, it wasn't for me. So I applied to medical school from that position, I got in, and then I really hated medical school. But it was the third thing I'd tried, and it felt too late to do anything else, so I just stuck it out.

The thing I hated was that most fields have these big organizing ideas, and medicine just has a lot of facts. You're memorizing a lot of facts, and I found it really hard to memorize them and to know why I was memorizing them. It was quite infantilizing — that's a strong word, but it was really just like being back in grade school. Then in third year — I thought it would be better, because I was in the hospital, closer to the reality of medical practice — but there too it was hard, because it felt like the only thing you were supposed to do was impress people. The goal wasn't really to learn; it was only to learn to the extent that you could impress the people you had to impress. So I had a fairly miserable four years in medical school, and that was mostly a reflection of me, not of medical school.

Then I took a fifth year off and just did research for a year, with a guy called Chris Murray, who was one of the authors of the first Global Burden of Disease study — an effort to systematize all of our knowledge of what makes people sick around the world, and use that to prioritize. That was a big idea, the first one I encountered that I could really sink my teeth into. I learned a lot about how to do research, and I came away pretty reinvigorated about medicine.

So then I started emergency medicine residency, and sadly I really didn't like the first couple of years of that either. Partly it was that I hadn't thrown myself into it properly. I wasn't a very good intern — I was one of those interns I now hate, because I thought I knew better than everyone else, and I hadn't read enough or seen enough cases of a thing to know what I was talking about. So I was kind of arrogant and just not very good, probably very frustrating to all the attendings I worked with.

Then at some point during my third year of residency, I started to realize I needed to get much better at this job, because it was really important. In your third year — even if, as in my case, you weren't particularly good — there are only a certain number of third-year residents, and there are shifts that need to be worked, so you get put into a position of responsibility whether or not you're ready. There was something about that that made it click for me: medicine was just really important.

So I focused on getting really, really good at the job, and the satisfaction I got from being good at it made me much more motivated — both to be even better, and to study it and try to make medicine better. At the same time as I was becoming a better doctor, I got really interested in some of the open clinical questions. In the emergency department — the A&E, sorry — maybe more than anywhere else in the hospital, you're constantly confronted by everything you don't know. The majority of people we see in the emergency setting, we send home or on to some other part of the hospital without ever really knowing what's going on with them. We can tell them what it's not — it's not a heart attack, it's not a stroke — but as far as telling them what they have, it's shocking how little we know about what's going on in people's bodies. Getting an appreciation for both the importance of this job and how much we don't know — this huge iceberg submerged under the surface of medicine — is really the thing that drives all of my research today.

Musty6:31

You said you didn't particularly enjoy medical school, and you didn't enjoy the first couple of years of residency either. So why carry on for six years if you weren't having a good time?

Ziad

Sort of desperation — the worry that I just wasn't going to like anything, that I was getting old already, and that I just had to carry on and get it done. And there were lots of things I could do with a medical degree even if I didn't like the training. So it was really incredibly bad decision-making. Had someone actually sat me down, elicited what I thought, and given me good advice at that point, I probably would have switched. It was very much a sunk-cost situation: I said, well, I've already put in all this time, I might as well get it over with. It's something I struggle with, because had I been making what I think would have been better decisions, I probably wouldn't be doing the stuff I'm doing today. So maybe it all worked out.

Musty

Could you tell me about the next step of the story, and particularly how your academic career came to be?

Ziad

After residency I took a job at the same hospital where I trained. After training, the experience of practicing medicine is completely different. Even when you're a senior resident, there's always someone else you can pass responsibility onto. When you're the one the responsibility ends with, it completely changes your decision-making. I started to get even more stressed about seemingly simple decisions, like sending someone home. I'd wake up the morning after a shift thinking about some middle-aged woman I'd sent home — "oh my god, I forgot to check this thing" — and I'd start calling patients.

I'd write my notes the day after my shifts, and as I was writing them I'd think, "I wonder what happened to this person, I hope they're okay," and I'd start calling people. I'd call five or six patients I'd sent home and check in. It was super interesting. And then I started wondering: it's really shocking that we don't know what happens to people after they go home. One of the really striking things is that if someone goes home and just drops dead, you might never find out about it. In the US there's a problem you don't have — the data are so fragmented that the death data aren't in the same place as the electronic health record data. So there's already this huge obstacle. But even in countries where the data are less fragmented, you might never know. It's only if that person happens to come back into your hospital, and one of your colleagues sees them, and a few days later says, "Hey, do you remember that patient?" And whenever someone says "do you remember that patient," you're like, oh, fuck — because it's never "yeah, everything went exactly the way you thought and they're doing great." It's always something not good happened, or you missed something.

So these are the kinds of things that started really stressing me out. And like any reasonable person, to deal with my stress I started doing research on the subject of my stress. I started a multi-year project on people who die after being sent home from the emergency department.

I was very lucky, because in the US the National Institutes of Health have a grants program for early-career people where you basically get five years and a lot of salary support to do whatever you want. I probably shouldn't describe it like that — it's a very structured program of research that nonetheless gives you a lot of liberty to pursue directions that are high-risk, high-reward. In fact, that's the official moniker: it's a high-risk, high-reward program. So I got one of those grants, and it gave me five years to think not just about interesting topics, but also to learn. One thing I learned about myself was that I really needed to know more math. I was lucky to start doing research with a few very good economists who mentored me and taught me a lot, and at the same time I did a lot of self-study to teach myself the methodological tools I felt I'd never gotten in medical school. That period was incredibly important — it showed me how much interesting stuff there was to do in medical research, and gave me time to learn all these methodological skills.

Musty

And then what happened?

Ziad

It turns out a lot of the things I was very worried about as a young attending — the things I was getting wrong disproportionately — were actually things that machine learning does very well.

In the ER, you see a patient and a lot of data about that patient, and you need to decide: what's the probability this patient has X, where X is heart attack, stroke, pulmonary embolism? You're taking in a huge set of information about a person and turning it into a set of probability judgments about what they have. Perfect job for an algorithm. Even my original research project — who's going to drop dead in the few days after I send them home from the ED — is a great machine learning problem. So a lot of the things I was most interested in and most worried about were things that humans do poorly and algorithms do very well.

I started getting really interested in how algorithms might be useful not just for helping decision-making, but for helping us understand better what was going on with these people. A lot of medicine moves forward like this: first you identify a group of people likely to have a bad outcome, then you study them, then you figure out what's going on. But first you have to identify them.

I have a project now on sudden cardiac death. Just to step back — sudden cardiac death is a huge unsolved problem. It's solved in one sense: we have a cure, which is putting in a defibrillator. It's unsolved in the sense that we don't know whom to put the defibrillator in. That's why we put in a lot of defibrillators that never fire, or misfire — by most estimates the majority don't fire — and still a lot of people drop dead.

So how have we made headway? Think about something like Brugada syndrome, one of the many things that cause people to drop dead suddenly. First a cardiologist noticed a link between a weird squiggle on an ECG and a bad outcome — a young person dropping dead. Then that cardiologist collected a bunch of cases and confirmed the pattern was there. Then the whole toolkit of modern medicine was deployed to figure out what was going on, and now we understand the genetic basis. When we see that ECG pattern, we immediately refer them for an electrophysiological study and to have a defibrillator put in. That's a great template of something we do as humans: we notice, we correlate, and then we study.

And that's an example of a part of the pipeline where machine learning can really help. One of the projects I'm doing now — in collaboration with a Swedish cardiologist — we've got all of the electrocardiograms done in an entire region of Sweden, and we're linking the waveforms to who drops dead of sudden cardiac death, thanks to the death certificates linked to all their ECGs and electronic health record data.

There are a lot of things like that where machine learning is going to be transformative, because it lets us deal with kinds of data we haven't been able to deal with before. How would you analyze an electrocardiogram using a traditional statistical toolkit? You can't. You have to code it with a human — the human has to say, "oh, the ST segment is elevated," and calculate the intervals. The only way we've historically handled that data is by putting it through the filter of human perception and understanding. But if that perception and understanding were complete, then nobody would drop dead, because we'd already know all the things that cause people to drop dead. So that's one example of the kind of task where machine learning is really going to be transformative.

Musty

Earlier you made a comment in that roadmap — first you identify a problem, or a group of patients who have a problem, and then work out what you can do next. That's a very problem-focused approach. But my feeling is — and I don't know if you think this is fair — that in ML in medicine, sometimes it's more of a solution-based approach, in the sense that you think, "here's this cool technology we have, now who can we use it on?" What do you think of that? Does it come with problems, or is it okay?

Ziad

Oh yeah, that's a great point. We focus a lot on "I have a new architecture for a convnet, what can I use it for?" rather than starting with "here's a really important problem that lots of people would like the answer to — how can I build for that?" Absolutely right. It's a big problem in the field, and it's why, for all the hype around machine learning — and I do think there's genuine promise — if you look at the actual examples of things that are genuinely useful, there are surprisingly few.

One way you see this: there was a really great paper that came out a couple of weeks ago in Nature, a system for targeting COVID testing at Greek border crossings. It's one of my favorite papers in recent memory, because it solves a real problem. There's a limited supply of tests; you need to figure out whom to test; and this way of thinking about the problem can really drive the efficiency of testing. They got 2x the efficiency they would have gotten from just randomly testing people. And if you look at the algorithm itself, there are a bunch of interesting problems they had to solve — related to the fact that tests happen in batches, and there's a delay between when you recommend the test and when you learn the result to feed back into your next round of learning. So real-world applications come with their own really interesting technical challenges.

It speaks to what I think the future of this field is going to be: applying it to real problems, and then using those real problems to push the methods further — not the methods pushing forward the problems. There's this general sense of how science progresses — that people sit at their desk, have an idea, and execute it. And the way we think about the scientific method is similar: I have a hypothesis out of nowhere, I go test it, I confirm or refute it. In fact, most discoveries are pushed forward by people trying to solve real-world problems. How has most of medicine progressed? In war. Why? Because there were a lot of problems that needed to be solved. How did physics progress? The same way.

So there are a ton of really important problems that will bring up their own technical challenges. There are these reciprocal links between the technical challenges and the real-world problems — but they need to start with the real-world problems. Otherwise you'll have a really elegant, beautiful paper that will never do anyone any good.

Musty20:26

To speak about medical AI being used on another real-world problem — would you mind giving a summary and talking about the story behind your paper that came out in Nature Medicine, "An algorithmic approach to reducing unexplained pain disparities in underserved populations"? I thought it was so interesting.

“So we trained an algorithm not to learn from the radiologist, but to listen to the patient.”

Ziad

Ziad

Oh, thank you so much. It was such a wonderful paper to work on. That work was led by Emma Pierson, a new faculty member at Cornell Tech in New York, and she's fantastic — a great example of someone trained as a computer scientist who really wants to use that training to solve real problems in a variety of fields, including medicine.

Let me tell you how the project started. I saw one of our co-authors on the paper, David Cutler, an economist, present a paper about pain. Pain is another really important problem for machine learning, because — how do we study pain? We try to correlate it to things on an x-ray or an MRI that might be causing it. Great machine learning problem.

With that background, I saw David present what I think is a medical mystery, and an example of the kind of real problem that demands a solution. The mystery is that when you look at a lot of surveys, there are huge disparities in pain — by social deprivation, by race, by education — and they're shocking. If you ask people whether they've been in severe pain in the past couple of days, in the US Black patients respond that they're twice as likely to have been in severe pain. These are big disparities, and it's very important — in this country and others it's a big driver of the opiate epidemic, which is just about pain.

What David presented was this mystery: even after you take into account the extent of, for example, osteoarthritis on someone's knee x-ray, non-white patients still report more pain. That's interesting, because it makes people think — where is that pain coming from? Their knee looks the same, so the pain has to be coming from somewhere else. Where this literature has largely gone is to focus on psychosomatic or psychiatric explanations. We know stress can cause the same physical stimulus to be rated as more painful — that's one mechanism. Another is that anxiety and depression are often expressed as pain, and the rate at which that happens can differ across linguistic and cultural backgrounds. Alternatively, doctors might be undertreating pain. Lots and lots of explanations for this important finding — but notice that none of them involve the knee. They all involve something else, somewhere else.

So I talked to David after that conference, and I said, "I bet it's the knee, and we're just not seeing it." He did not share my view. But we decided to actually look into it. In some ways this paper was settling a bet that David and I made many years ago.

What we did was try to train an algorithm to help settle the bet. But the way you'd normally train such an algorithm wouldn't help you settle it. The way the machine learning playbook is typically applied to radiology is: take the x-ray and output the degree of arthritis on it. But how do you measure that? You ask a radiologist. If my position was that the radiologist is missing important causes of pain that affect non-white patients more, then that approach isn't going to help us find those causes — because the algorithm is just learning what the radiologist says.

So there's a bit of an impasse, because all the data available to us thanks to electronic health records is the x-ray image paired with what the radiologist said about it. A big obstacle is the lack of data labeled with real outcomes, as opposed to human judgment, that we can use to train algorithms that learn from nature, not from humans.

We were very lucky to find a dataset that had the x-ray images of a lot of patients from a very diverse sample, paired not with what the radiologist said about the knee, but with what another important human said about the knee — the owner of the knee. The person whose knee it was, rating the degree of pain that knee gave them day to day. So we trained an algorithm not to learn from the radiologist, but to listen to the patient — to guess, given an image of a knee, whether this knee is going to be painful or not. It turned out that was quite different from the rating the radiologist gave. We found a lot of knees that looked much more painful than you'd have thought from just looking at the radiologist's report. That echoes what we know from a lot of the literature: there's only a moderate correlation between the degree of pain a patient reports and the degree of arthritis a radiologist sees on the knee.

But the particularly interesting fact was that that disparity was greatest for non-white patients — and, to a somewhat lesser but still present extent, for lower-income and lower-education patients. The algorithm was finding things on the x-ray that the radiologist wasn't seeing, and that was disproportionately correlated to increased pain in disadvantaged groups.

It made us really think about what could be going on. One hypothesis: how do radiologists learn to judge the degree of arthritis on a knee? They learn it in medical school and residency, and that knowledge goes all the way back to studies done in Lancashire in the 1940s and '50s, looking at the knees of coal miners and comparing them to the knees of office workers — figuring out the radiological appearance of knees that were linked to complaints of pain. When you go back to those studies, you don't even find, in the methods section, a description of the racial or gender breakdown of the population — because it was all the same. It was all white male coal miners, which was very reasonable at the time; that's who they were studying. But our medical knowledge has never caught up with the fact that we see patients today who are not white male coal miners. Even though that knowledge does a surprisingly good job of extending beyond white male coal miners, it's obviously missing some stuff too — and that's the stuff the algorithm was able to see in those x-rays.

So, to go back to our initial mystery — this pain gap between Black and white patients, where you take people whose x-rays look the same to the radiologist and Black patients report much more pain — when we substitute the algorithm's judgment for the radiologist's, we were able to cut that unexplained pain gap by nearly half. Even in this fairly small sample, the algorithm was able to dramatically explain the unexplained pain in these patients. We're hoping to scale that study up by bringing more data online so other people can also do this work — and that's through a nonprofit I co-founded with one of my co-authors, called Nightingale Open Science, which we're launching at NeurIPS in December.

Musty

What's so amazing about that story is that if someone was just taking the solution-based approach — apply the technology to whatever problem it fits — the first thing they'd have done is take this dataset and train an algorithm on what the radiologist thought of the severity, and that would come with all the inherent biases and systemic issues. But by taking the problem-based approach and thinking, "how do we actually solve this problem," you sidestepped that and used AI to remove bias — which I haven't heard of being done very much. It's usually the other way around, right? AI systems come in and increase bias.

Ziad

Yeah, you're totally right to put your finger on the problem-based approach. We went into that paper trying to solve a problem — trying to understand whether there was something radiologists were missing that could explain this mystery in the literature. And what is the problem we're trying to solve when we're just automating the radiologist's judgment? It's not totally clear to me. It could lead to automating radiologists — but if you know anything about automation in other fields, that's not the way automation generally works. These tools can be incredibly powerful when they're turned to answer interesting questions, but when they're just applied expediently — because you've got an algorithm that works well on detecting cats, and now you've got another dataset of x-rays, and your goal is just to show that your cat detector is also a pneumonia detector — it's not going to produce the same level of interesting results.

Musty31:42

Another thing I got from reading that study was this frame shift: instead of trying to emulate the best doctors or radiologists and get a result as good as them — them being your gold standard — why don't we try to supersede, and become almost superhuman? Are you seeing other low-hanging fruit for that kind of approach — places where we can actually become better than humans, rather than aiming to be as good as the best human?

“We're setting the bar very low if all we want from algorithms is to replicate our own judgments with all of our own errors and biases.”

Ziad

Ziad

Yeah, you put it exactly right. We're setting the bar very low if all we want from algorithms is to replicate our own judgments with all of our own errors and biases. That would be a very depressing future world to aim for. The goal is very much to have algorithms that don't just learn from humans, but that learn from nature — from patient outcomes, from patient experiences, from all these things we want to explain but can't.

The knee pain paper is one example. We have another paper I'll mention, because it's another great example — helping emergency physicians test for heart attack better. We trained an algorithm to find people coming through the emergency department who are having an acute coronary syndrome. There's one very striking result. In the paper we compare the complexity of the model the algorithm uses to predict heart attack — how many variables, how much structure — with the complexity of the model physicians are using when they decide whom to test with a stress test or a catheterization.

There are two interesting facts. One is that physicians, at least in our frame, are using a handful of variables to make their testing decisions — you can explain a lot of the variation in who gets tested with a small number of variables. And for those variables, physicians are using them shockingly well. The physician's model is actually pretty close to an optimal machine learning model of that size. If you told the machine learning model it could only use 40 variables, it would actually look a lot like the physician's model, which is really amazing. So physicians are doing a really good job with the variables they can hold in their heads and process.

But the problem is that heart attack is not a handful of variables. Heart attack is super complicated — it's this mystery of why a thrombus develops at a certain time in a certain person. It's very clearly not a simple phenomenon. The reason the algorithm does so much better than the physician is that the algorithm can hold not 40 variables but 40,000 variables in its memory, process them, and use all of them to form a well-calibrated probability judgment on heart attack in a given person at a given time. So it's not a fair competition — when we have these kinds of data, algorithms can do much, much better than humans at allocating testing.

We take advantage of a little natural experiment built into our data. Depending on the exact moment a patient arrives in the emergency department, they're going to be triaged by a certain team at the triage desk, and those triage teams have a higher or lower likelihood of sending patients on to a stress test or catheterization. When you look at the high-testing shifts versus the low-testing shifts for everyone, we can't really detect much of an effect of testing more. That's the standard "less is more" story — you scale up testing, you don't find any more heart attacks, you're just wasting money. But when we hone in on the patients the algorithm predicts to be at very high risk of heart attack, and those patients get assigned to a high-testing shift instead of a low-testing shift, their one-year mortality is about 30% lower.

So it goes to show that it's true that less is more on average — but medicine isn't about averages, it's about individual patients and what they need. Machine learning can really help figure out which patients need what. In this case, by figuring out who needs to be tested, we can cut testing for a lot of the low-value patients doctors are currently testing, and actually scale up testing in the small fraction of patients doctors are currently not doing a great job of testing — the ones who need it at 100%.

Musty36:43

This is a really broad question, but — you mentioned that when a doctor looks at someone with chest pain they might be weighing up 40 variables, and when the algorithm looks at someone it might be considering up to 40,000. In machine learning generally, and in healthcare specifically — for someone like me who doesn't know much about the field, I might think, okay, if this is using 40,000 variables instead of 40, it's going to be exponentially better. More variables is better, more data is better. Is that something you generally see, or is it not really the case?

Ziad

On average, the more data you have the better — and that goes for both the number of observations (how many patients are in your dataset) and how many variables you measure on those patients. But one of the really striking things is that 400 variables is not 10 times better than 40 variables. It's probably like 2 or 3% better. So you really need a lot of data.

One of the interesting contrasts to humans is that we do a shockingly good job of figuring out patterns with very little data. Think about an algorithm trained to identify a cat and distinguish it from a dog — that algorithm needs a lot of cat pictures to figure out what a cat is. But you show a child a cat, and the child knows what a cat is; you show that child a second cat and the child says "cat." You don't need to show the child a picture of the cat upside down and sideways and do all these things we have to do for algorithms to make sure they understand what a cat is. It's a really striking thing, how much data these algorithms need to figure out what a heart attack looks like versus how much data a human needs. It's one of the mysteries of human intelligence that I've been thinking about a lot, as I see in some ways how limited algorithms are and how much data they need to learn the same kinds of things.

That said, we do have those datasets, and we do measure a ton of variables on a ton of people. So bringing more of those datasets online is super valuable. I think that's the single biggest problem in this field right now — the lack of availability of good data for people to train algorithms on. That's the goal of that nonprofit I mentioned, Nightingale Open Science. It's working with health systems to create interesting datasets — largely focusing on imaging and waveforms — to label them with ground-truth outcomes like patient outcomes and patient experiences of pain, de-identify them, put them on our cloud, and make them available to researchers free of charge.

Musty40:31

When we look at other fields where algorithms and AI are used — banking, insurance, welfare systems — there's been more widespread adoption and normalisation, and we're seeing some of the problems these can cause for disadvantaged groups. Up until now these haven't been as widely used in medicine, or implemented in day-to-day clinical practice. You might disagree with me here, but in general, what's your feeling about this move towards these kinds of systems being used more widely in medicine? Are you cautiously optimistic, or do you think this is going to potentially go wrong?

Ziad

It is terrifying. Even though you're absolutely right that these kinds of automated decision-making systems are not widespread in the clinic, there are many other parts of the hospital and the health system where they are very widespread.

We have a paper we published a couple of years ago in Science that looked at the class of algorithms used to make population health management decisions for patients. This is essentially trying to identify patients with health needs that are emerging over time, so we can help them today. The stereotype would be someone with congestive heart failure who's on a slightly too low dose of their diuretic, so they're slowly deteriorating and end up in the emergency department — and all of that could have been avoided had we just known about the creeping fluid overload and gotten them on the right medication. There are lots of situations like that, with diabetes and so on, where we want to help those people. We have resources to help some of them but not all of them — it's a scarce resource — and we want to get the right people access to it.

We studied one algorithm, made by one company, that in the US alone is being used to make that decision for 70 million people every year. And that family of algorithms, which all work the same way, is being used — by industry estimates — for 150 to 200 million people a year. So the majority of the US population. The scale has already gotten enormous, bigger than in many other systems like criminal justice or finance, where largely people are still making different kinds of decisions that are not algorithmic.

We found that that algorithm, and that family of algorithms, had an enormous amount of racial bias. The cause was that the algorithm developers — not just at one company, but at companies, hospitals, academic groups, and parts of the federal government — had decided to measure someone's health needs in terms of how many healthcare costs they generated. Which is not unreasonable, because people who get sick do generate costs. But unfortunately it's a biased measure of health needs, because not everybody who needs healthcare gets healthcare.

And this isn't just about the insured — the sample was well-insured, to the point that they'd be very comparable to European national health insurance schemes. So it wasn't about insurance. It's just the fact that not everybody who needs healthcare gets it, and that's not equally distributed in society. The most vulnerable people are least likely to get healthcare even when they have insurance — because of transportation, or getting a day off work, or miscommunication, or distrust of their doctor. Given the scale and the size of the bias we found, everyone should be very concerned about that in healthcare, just as we should be in criminal justice and finance and everywhere else.

But that work has left me cautiously optimistic, in the sense that the algorithm can be fixed. We actually worked with the company that made it to create a new version, trained to predict not someone's healthcare costs but measures of their healthcare needs. When we did that, we dramatically reduced that bias.

Once you know what you're looking for, it can be fixed — and when you fix it, it turns the algorithm from a tool that reinforces structural inequalities into one that actually fights against them, by getting resources to people who need them rather than people who already have them. So yeah — cautiously optimistic is a good summary.

Musty

Have there been any habits or ways of approaching things that have been helpful for you?

“The impact you can have with one good article is just huge — much bigger than the impact you can have with lots and lots of short articles.”

Ziad

Ziad

One thing I've been trying to do recently is dramatically cut down on the number of things I'm working on. I think that I — and pretty much everyone I know — am working on too many things. To some extent that's a distortion imposed by the fact that we publish, and we value a large number of not-good articles rather than a small number of very good ones. I think that's too bad.

But even with that constraint, I've found that the impact you can have with one good article is just huge — much bigger than the impact you can have with lots and lots of short articles that would take you the same amount of time to produce. So I'm trying to really focus in on a handful of projects that will actually lead to some different state of the world — if what I think about that paper is true, and it gets published, and people pay attention. It's really easy to get distracted and pulled in lots of different directions. It's been both useful and much more fun for me to spend more time thinking about a smaller number of ideas and pushing them forward over time.