Mission 11 // May 23, 2020

Validation over Valuation: An Oncologist's Battle Against Babylon Health

An oncologist's three-year fight to prove Babylon's AI chatbot was missing heart attacks — and why no regulator acted.

DW David WatkinsConsultant Oncologist, The Royal Marsden NHS Foundation Trust
Validation over Valuation: An Oncologist's Battle Against Babylon Health
0:00 // 17 min

About this episode

For those of you not familiar with Babylon, it's a UK medtech company endorsed by Matt Hancock and recently valued at $2 billion. It has two relevant parts. On one hand it's a telemedicine service which, although it has been criticised, is nowhere near as controversial as the other half: the Babylon chatbot. The idea is that a patient reports their symptoms by playing a game of 20 questions with the chatbot, which then uses Bayesian reasoning to suggest some potential causes. Crucially — at least from a regulatory standpoint — Babylon says this isn't offering a diagnosis; it's simply an information-giving exercise.

Dr David Watkins, a consultant oncologist who was going by "Dr Murphy 11" at the time, started testing the chatbot and noticed some unusual results, which he began to share on Twitter. This resulted in what I think can safely be described as a feud between Dr Watkins and Babylon — one that culminated in a fiery debate at the Royal Society of Medicine, and then an appearance on Newsnight by Dr Watkins and Babylon's Dr Keith Grimes. I started off by asking Dr Watkins if he could catch us up with the story.

In this conversation

  • A three-year clinical detective story: how a single consultant oncologist, tweeting anonymously as @DrMurphy11, kept catching Babylon's triage chatbot advising heart-attack symptoms to wait for a GP appointment.
  • The example that went viral — identical mid-60s smoker with chest pain and nausea: entered as a man, it flags cardiac causes and sends you to the ER; entered as a woman, it says "panic attack, deal with it at home."
  • Why Babylon's published "evidence" doesn't count: they wrote their own exam, sat it under their own conditions, marked it themselves — preclinical testing with no real patients — then claimed 100% safe.
  • The counter-intuitive red flag: zero reported adverse events across 4 million users isn't reassurance, it's a sign nobody is looking for the incidents that must be happening.
  • The thesis in the title — technology should earn trust on its value and evidence, "not on the basis of whether they're mates with Matt Hancock" — and what it took for a doctor to give up anonymity to say it.

Transcript AI-generated

“Babylon were making some very bold claims about the accuracy and safety of the chatbot. Around that time they were claiming it was 100% safe, and as accurate and as fast as a nurse or doctor providing triage advice.”

David

David

So I first came across Babylon in February 2017. It was something I noticed on Twitter, and at that time they were promoting the rollout of an NHS 111 system in North London that was going to be piloted — a chatbot triage system. I'd never come across a chatbot triage system in the NHS before, so I was interested. I downloaded their app and gave it a go.

Right from the outset I identified some very basic and fundamental flaws. Some of the questions and responses you'd get from the chatbot were, quite frankly, absurd, and I noticed some safety concerns. There was a scenario I tried very early on: a patient presenting with classic symptoms of a heart attack — central chest pain — and the chatbot advised them to get a GP appointment sometime later that day. It didn't say to call an ambulance or go to A&E. So I flagged those concerns to the CQC, the Care Quality Commission, and I expected them to deal with the patient safety concerns.

It also became apparent that Babylon were making some very bold claims about the accuracy and safety of the chatbot. Around that time they were claiming it was 100% safe, and as accurate and as fast as a nurse or doctor providing triage advice. Over the next 16 months or so I continued to flag concerns and query the CQC about what action they were taking, through to the summer of 2018. At that point I got a response from the CQC saying, essentially, not our problem. They didn't take any action to address the safety concerns, which was disappointing. And at that same time Babylon did another bold promotional event at the Royal College of Physicians — they had the chair of NHS England in attendance — where they claimed their chatbot was as good as a doctor at making diagnoses.

The MHRA became involved at that time, because the safety concerns were still there. These flaws in the algorithms were just as dangerous in the summer of 2018 as they'd been in early 2017 — no action had been taken. Unfortunately, due to a variety of loopholes and limitations in the regulatory system, the MHRA have been largely ineffectual in addressing the safety concerns. They have been proactive, they have been engaging with me. Throughout 2018 and 2019 I raised concerns through them — I think I raised 25 queries in February 2019 alone. About six months later, in August 2019, I got responses back, retested the system, and found the flaws were still there.

So it was apparent that Babylon weren't taking the safety concerns seriously, and the regulators either weren't taking them seriously — from the CQC's perspective — or weren't able to take definitive action, from the MHRA's. At that point, in August 2019, I decided that as an anonymous individual I didn't really have a strong enough voice to air these concerns. I made the decision to go public, and was then invited to speak at an event at the Royal Society of Medicine in February of this year. That tied in with the piece Newsnight covered on the same day. And here I am, speaking to you now.

Musty5:44

What was Babylon's response to you raising the concerns?

David

That's one of the disappointing aspects of this. These concerns were raised in good faith, initially on Twitter, to Babylon — and there was an absence of response over a month or so. At that point it was suggested that they should be raised with the CQC. And for a prolonged period, Babylon just did not engage. Between February 2017 and June 2018 there's no engagement from Babylon to speak of.

When the concerns did come out in public, instead of saying, "OK, yeah, we're aware of these issues, they've been flagged to us by the MHRA, we're addressing them, we'll do our best to ensure patient safety is paramount" — they didn't. They attempted to discredit the person raising the concerns. That was very disappointing. The concerns went public in an article in the Health Service Journal in June 2018, which came about because I shared my correspondence with the CQC and MHRA with them. Babylon's public response was to suggest it was an anonymous critic making misleading claims because of vested interests — none of which was true. The incidents were raised in conjunction with the CQC and the MHRA.

Musty

Can you talk to me a bit more about the red-flag issues — what kind of things you'd put into the chatbot, and what kind of results you'd get?

“If you put that in as a man, you were directed to go to A&E. If you put it in as a woman, it said you're having a panic attack — don't worry about it, you can deal with it at home.”

David

David

The persisting issues over the years have been chest pain — and probably one of the best-known ones was the panic attack. Quite often, when you put chest pain into the chatbot, it would come up with the suggestion that you weren't having a heart attack, you were having a panic attack. Right through to November 2019, that persisted. At one point someone on Twitter said, is it the same if you do it as a female versus a male? So I tried it both ways.

The patient was a mid-60s smoker who develops chest pain and nausea. If you put that in as a man, you got a suggestion of cardiac causes and were directed to go to A&E. If you put it in as a woman, it said you're having a panic attack — don't worry about it, you can deal with it at home. And the fascinating thing is, anyone with an ounce of sense would think: hang on, just because you're a woman doesn't mean you can rule out a cardiac cause for chest pain.

Babylon, again, instead of addressing the concern, dismissed it. They said the data supports this — women are more likely to have panic attacks and less likely to have heart attacks. I find it remarkable that instead of addressing these issues they try to justify the flaws in the system. And that's what makes the system so dangerous.

Musty9:32

Babylon have produced some evidence for the chatbot — they've posted a couple of trials on their website comparing it to medical doctors. What do those trials show, and what do they not show?

David

The issue with the data Babylon have put in the public domain is that it's essentially preclinical testing — the sort of testing you do before you actually go out and test in real patients. No patients were involved. What they essentially did was develop their own exam paper, undertake it under their own conditions, mark it themselves, give themselves a grade, and say: we're fantastic, we're 100% safe. There's a complete lack of independence.

And there's no real-life data. This did not involve patients in any way. It's very much the first hurdle in the development of an e-health technology or any medical device: you do your preclinical testing, then you go and do a study in patients. You evaluate your device, you validate its safety and accuracy, and through that you show the device can be trusted to give patients appropriate advice. But they've never done that. They've had a good few years now, so there's been plenty of opportunity — and they've failed to, for whatever reason.

Musty

And from what I read of the trials, the implication was that their chatbot is equivalent to a human doctor.

David

If you've done those studies and you make appropriate claims on the basis of them, then fine. But you shouldn't make bold claims that aren't supported by the evidence you have. They've never tested the chatbot in real-life settings against regular doctors, so I don't think you can make those claims.

Musty

Any AI or chatbot is going to have its faults — as do human doctors. Babylon's claim was that you ran 2,400 tests, and only around 100 of them were raised as concerns, which would suggest quite a low error rate. And to their credit, they say they've never had a single adverse event reported due to the chatbot.

David

It's interesting that you say "to their credit, they've never had an adverse event." From my perspective, I perceive that as a very serious red flag.

With regards to what Babylon say — I think it's become clearly apparent that Babylon cannot be trusted, so you can't take their word at face value. I can tell you there were nowhere near 2,400 tests undertaken of different triage scenarios. At different times the chatbot has varied in its degree of accuracy and safety. But to take one example: one of the concerns I raised right at the beginning was this risk of a missed heart attack, where patients presenting with cardiac symptoms weren't being advised to go to A&E — the chatbot would say it's just heartburn, or a panic attack, and so on. Over the three-year period I flagged that on Twitter on 28 separate occasions, between 2017 and the beginning of this year. And throughout that period Babylon failed to address that single flaw — that very basic, fundamental flaw, which should never have existed in the first place. So one of the other claims Babylon made — that they fixed every issue almost immediately — it's just not true.

As for the fact that they haven't reported any incidents: this is a chatbot apparently in use by 4 million people around the world. They claim that when patients interact with it, 40% of the time they simply rely on the advice of the chatbot and don't seek further medical guidance. If that's the case, there should be hundreds, thousands of incidents being raised — because if you're providing that much medical advice, you get things wrong. No technology is 100% safe. So if they're not getting these reports, it means they're not looking for them, or people aren't sending them for whatever reason. If you're looking after that many people, you should be getting a hell of a lot of incident reports. Go to any hospital or GP practice and they'll be flagging patient incidents on a weekly, daily basis. So if Babylon aren't, there's an issue with their vigilance.

Musty15:12

I hope you don't take this question badly — but what would you say to someone who compares your plight to that of the Luddites?

“It's not about bashing innovation — it's about celebrating true innovation.”

David

David

I don't take that badly. It's not about bashing innovation — it's about celebrating true innovation. We have to be confident. Members of the public want to be confident, and we as healthcare professionals want to be confident, that technology is used on the basis of its value, its worth, its evidence — not on the basis of whether they're mates with Matt Hancock. This isn't tech bashing at all.

There are so many alternative systems out there which haven't had a look-in over the past two years, because the narrative has been Babylon, Babylon, Babylon. That's what's frustrated me. And why is it me? It's a fascinating question — it really shouldn't be me at all. But I think it flags an issue within the tech sector: there's lots of overblown hype and PR involved, and sadly, often not enough evidence to back up the claims being made.

Musty

I hope you enjoyed that episode. You can find Dr Watkins on Twitter at @DrMurphy11, and you can find me at www.bigpicturemedicine.co.uk.