Mission 133 // October 27, 2024

A New Radiology Foundation Model ($160M Raised)

Harrison.ai's models read millions of scans a year. Aengus Tran on building clinical-grade AI from Sydney and taking it global.

AT Dr Aengus TranCEO & Co-founder, Harrison.ai
A New Radiology Foundation Model ($160M Raised)
0:00 // 35 min

About this episode

Draft page — this summary was written from the archive and should be checked against the released episode before publishing.

Dr Aengus Tran is the CEO of Harrison.ai, an Australian radiology startup that has raised $160 million in funding. He co-founded Harrison.ai with his brother, both having grown up in Ho Chi Minh City before moving to Australia.

They are used by 50% of radiologists in Australia and will now triage one third of all chest x-rays in NHS England. The company's current radiology solution is cleared for clinical use in 40 countries, and they've just launched their radiology foundation model, Harrison.rad.1. Here's what's crazy: when sitting the FRCR 2B radiology exams, it performed on par with human radiologists and demolished existing models like GPT-4, Claude and Gemini — scoring around 51.4 out of 60, whilst the others scored about 30 out of 60. By the way, these are pretty tough exams. About half of human radiologists fail them the first time around. Pretty impressive.

By the way, when you live on a tiny island like the UK or Australia, you dream about US domination in your go-to-market strategy. But for Aengus, being on a tiny island was kind of a superpower.

In this conversation

  • Harrison.ai's radiology foundation model, Harrison.rad.1, scored ~51 out of 60 on the FRCR 2B — on par with human radiologists and roughly double what GPT-4, Claude and Gemini managed on an exam half of humans fail first time.
  • The "rest-of-world-first" bet: Aengus is glad he didn't start in the US, because the FDA only clears narrow AI — so the real automation (and the real value) is everywhere else first.
  • The secret behind superhuman accuracy: consensus labelling by 250 clinicians plus "time machine labelling" — showing labellers the CT scan taken years after the x-ray, so the model learns to predict cancer before a human ever could.
  • His defensibility framework: the model is just one layer of the "innovation stack" — product, medical-device clearance and reimbursement sit on top, so if GPT-8 ever wins, he'll just swap in their model and keep the rest.
  • A masterclass in scrappiness at healthcare-grade stakes: flying to Vietnam to build the region's biggest radiologist labelling operation, and buying a gaming-PC GPU cluster every time they raise a round.

Transcript AI-generated

Musty

You did something quite smart, which was that you have Harrison.ai and then, very quickly, you formed this joint venture with I-MED Radiology, who have a stupendous dataset of radiology images. I don't know whether they're labelled or not, but they have a large dataset. That's quite a smart move, right? You immediately realised that was an important part of this journey, and you almost partnered with them straight away. That's my understanding of what you did.

Aengus

Yeah. In the early days of Harrison.ai, I had a previous startup — it was in the field of human fertility, actually. This is before Harrison. I built an AI in IVF that could select human embryos. This was all done while I was still a medical student at the University of New South Wales. It was a trade sale to a European-listed company in 2018. One of the lessons I learned from that was how important it is to have a clinical partner — a clinical operator — to help build the technology. And it's for a couple of reasons.

Number one: even though I'm a clinician, the workflow is very deep. A general clinician may not fully appreciate what it means to operate radiology, or IVF, or pathology. So having a clinical partner who shares the clinical concerns and the problem to be solved is so key. That's why it's always been a partnership model for Harrison.ai — to find operators of healthcare, service providers, and work with them on cracking the right solution.

The dataset is the other nice part of the partnerships — obviously a very key ingredient in building AI. And we want to make sure we do all of that ethically and responsibly, by anonymising the data and following the relevant guidelines and regulation in each geography.

The step beyond that is the deployment. I call it the Kickstarter model, but for medtech AI.

When you build technology in healthcare, the hardest thing is getting that initial adoption, because it's very hard to do an MVP in healthcare. No one wants an MVP. How would you like an MVP heart valve? You want the full thing, right? You want all the bells and whistles. You want it to work end to end, everything coming together. Which means that when you're incubating a new technology in healthcare, you can't just release something that half works and then iterate over time — like everyone else would tell you if you read startup books. You almost need a very clear vision of what a complete product looks like, guided by the people who'd actually use it, and then go to that minimum level as your first release. And that's part of the clinical partnership: we can do a very safe iteration with a partner to ensure that what we end up releasing is fit for purpose on day one.

Musty

I'm really curious, because the way you're describing things — they sound very methodical, well thought out, intentional, and generally quite professional. But startups aren't like that. They're kind of scrappy. It's kind of "let's just build some MVP, figure it out, and then next step, next step, next step." Can you tell me any stories from the early days where you had to embody that spirit? Because I know in healthcare you can't always have that MVP approach, or that scrappy approach you can have in other industries — but were there elements of that in your story?

Aengus

Yeah. One of the first things we had to figure out, after we decided we wanted to be in radiology and got the data sorted, was the labelling. If you cover healthcare or AI, you'll know that the other thing no one ever talks about is that big labs — OpenAI, Anthropic, even Tesla — pay companies like Scale AI billions of dollars to do data annotation. That is the secret behind the incremental performance gains we're seeing today. So one of the first things we needed to figure out was: how do we do data annotation at scale for medical diagnosis?

Unfortunately, we couldn't just hire anyone off the street to label a street sign, and we can't hire a PhD to solve a maths problem — medical data requires medical professionals to do the labelling. So in the early days, I remember us jumping on a plane to Vietnam. Dimitry and I grew up in Vietnam and have deep connections with the healthcare community there, so we set up a data refinement operation. We wanted it to be the biggest of its kind, so we did the math — we had an Excel sheet — and we'd need to hire something like 250 clinicians to do this. It would make us the biggest employer of radiologists in the region, or something like that. This company has become really big now, but at the time we were like, what do you mean, 250 clinicians to do annotation? We didn't even know where to begin.

But we started with a team of 10. We managed to convince the radiologists — they'd never even heard of AI — but we explained that this is basically teleradiology, except you don't have a patient on the other side you need to care about.

This is retrospective, anonymised data, and you perform the annotations to teach a machine system how to do the same task. It was very challenging in the early days, because quality control is very difficult. This is the manufacturing line for AI.

We pioneered a consensus-based process where each scan is labelled multiple times by random radiologists from the team, with various levels of experience. We have a system to automatically estimate their performance on individual labels and use that to weight their opinions against their peers. So it's a complex system — we almost had to reinvent how radiology works to get this data annotation up and running. But eventually we did get to the point where we had 250 clinicians doing data annotation. And that's one of the other secrets behind the performance of our system: it has learned not from a single clinician, but from a consensus of experts. That's why it's managed to outperform human — even expert — radiologists today.

Musty

I think it's pretty obvious in medicine why an AI or computer system would be superior to a human interpreting radiology images. Whether that's today or tomorrow, it's pretty intuitive why that would be the case for that kind of task — it just kind of makes sense. But what I'm really curious about is: if you make this delineation between AI and the human mind, there must be some things the human mind is better at than a machine. One that comes to my mind is that humans are actually pretty good — at least from what I've seen — at picking out edge cases. They might see one case once in their lifetime, twenty years ago, and they'll remember it for the rest of their life; they won't need it reinforced a hundred times. I'd love for you to opine on where you think the human mind is actually, today, superior to AI.

Aengus

I think what you're describing, in AI we call zero-shot learning. The human mind is very good at zero-shot learning — you give it one example. If you look at GPT, you can say: here's an example, here's another example, here's a task, please perform it. That's what we call zero-shot or few-shot prompting. Humans are really good at that. You can show them a single example of a character and they'll be able to identify that character in any pose, any costume.

The model would really struggle with that. So that part, we still haven't cracked — we need a lot of examples. It's very data-inefficient to train these models compared to a human radiologist. But that's actually the purpose behind Harrison.rad.1, the foundation model we've just built. Up to this point we've had a foundation model, but we've largely used it to fine-tune into higher-performance classifier systems. The foundation model we've built is meant to be more general-purpose.

One of the things we really hope to get it to do is this sort of zero-shot learning, where you can give it an x-ray and say, "This is the Mustafa finding, the Mustafa sign," and "here's the Aengus sign," and "here's another example of a Mustafa," "here's another Aengus sign" — you give it four or five examples, and then eventually it goes, okay, I get you now, I know what you mean, and I'm going for it. It doesn't work as well as the human right now, and we're still doing a lot of work to understand why. One of our researchers and clinical AI leaders, Jarrel, who's a radiologist on the team, is very passionate about this topic. Because once you crack that, you solve the last bit that separates the human from the system today: the ability to rapidly learn new things based on existing base knowledge.

Musty

Very interesting. Look, Aengus, feel free not to answer this if you don't want to, but I'm quite interested. I can imagine that running Harrison.rad.1 — let's just say running it to process one chest x-ray — is quite computationally intense and therefore expensive today. Are you able to share a rough sense of what it would cost? I don't know whether you'd call it one call, or one image run through, or one interpretation. Can you give me an idea?

“Medicine has a scale problem, not a science one.”

Aengus

Aengus14:33

Yeah. We're thinking about pricing that's very similar to the other frontier models from a pricing-model perspective — so price point per million tokens. We think it'll be a premium model because of its unique capability in healthcare. It's interesting, because the economics of a foundation model in healthcare are very different from a foundation model in general.

You can think about GPT-4o as a very well-educated intern performing intern tasks — that's the analogy a lot of people use, in coding or in email drafting. If you think about it, an intern doesn't get paid a lot. They get paid $50,000 or $60,000 a year. So ultimately the tasks these models are automating today are very low-skill, low-value. And there are many of them — there are hundreds of these models now. So the commoditisation pressure is driving the cost down dramatically. They're open-sourced as well. Eventually it converges to something a little bit higher than compute cost — you pay AWS and then make a tiny margin on top, if your model is significantly better.

In healthcare it's very different, because clinicians, radiologists, are some of the most well-paid people on the planet. That's the reason why, in Asian culture, all the kids get pushed to become a doctor, right? It's a bit of a joke, but it's not so much — because in Australia, an average radiologist could get paid up to a million dollars a year if you factor in bonuses. So the task we're trying to solve is much lower-volume but higher-value, if that makes sense. There are only a billion chest x-rays done in the world every year. We don't need to do that many inferences.

It's much more important that each of our inferences is extremely accurate — like healthcare-safety accurate. So for us, the cost to serve this model is very easy to wrap our heads around. There'll be far fewer of these models — not hundreds. There'll be a handful of models that are actually accurate, and the problem being solved drives more value. So the economics of a healthcare foundation model are much better than a best-man-speechwriter AI model, if that makes sense. At least that's how I think about it.

Musty

What I think is quite interesting — and I'd love for you to talk about this — is that in your entrepreneurial strategy, you seem to like this concept of taking big bets. The reason I think that, and I'd love for you to comment on whether you agree or disagree, is that even with Harrison you had this big bet: you were going to spend two years on R&D without any go-to-market, seemingly without any deployment, when that's not always the conventional advice. That could have backfired — you could have spent two years making junk with no deployment and no income. And now there's another critical point: you've made an investment into a radiology foundation model, which again is a big step-function investment in R&D, but with the goal that the whole world runs on your model and everyone needs to license your infrastructure layer of radiology. That's what you've done again. Can you talk a little about that? Do you agree — taking big bets, that kind of thing?

Aengus

Yeah. The two things you've just mentioned actually follow the same arc. When we started Harrison, we made the observation that there are 1.5 million radiologists and pathologists missing globally, and it takes 15 years to train one of them — six years of med school and nine years of subspecialty training. So we're 22.5 million years behind in terms of training years. It's an impossible problem — not solvable with training alone. I challenge anyone to put forward a credible plan to solve that capacity problem today. We fundamentally made a long-term bet that AI automation is the only viable solution to the capacity issue, in any timeframe, with any investment. I've yet to see another credible plan that can solve it.

Education is hard because medicine is an apprenticeship — which means that to train more doctors, you need to take more doctors out of the system to supervise them. So there's a natural limit on how many doctors you can train every year.

I don't think a lot of people realise that, but it's called an internship for a reason — because you're interning with someone. So if you take that view, then the thing we actually need to build is the ability to automate diagnostic tasks, and therefore shift the role of radiologists and pathologists from very mechanical interpretive tasks — reading an image and writing a report — into higher-skill tasks: performing procedures, managing edge and corner cases, procuring and configuring fleets of AI systems for their organisation. This is what the biochemistry department in every hospital already does today, with blood tests.

So that's the bet. That's the meta-bet, if you like. And everything along the way has been to move us in a straight line towards that. Of course, along the way we want to build valuable intermediaries — things that deliver real-world value. We don't want to take a five-year shot toward the moon. We want to go to outer space — but we want to do high atmospheres, we want to do rendezvous. There are multiple steps along the way we want to go through. As CEO of Harrison.ai, one of the big challenges for me is: how do we do that without taking a detour? What are the milestones along the way we can interpolate to, while not being distracted from the real objective — automation in diagnosis?

Musty

Okay. I want to do another thought experiment with you — I think this will be fun. Imagine you look into a crystal ball and you see that, let's say, GPT-8 will come out in five years' time. There'll be a step change, and it'll far outperform whatever you can build with your datasets and your own foundation models. It'll just commoditise image interpretation in radiology — that's a fact — and it'll be 100% performance, accuracy, etc. What would you do today if you knew that was the case? What would be your next moves?

“In the UK, a third of NHS England's chest x-rays run through our system. That's a journey we took, going with our partners at the NHS.”

Aengus

Aengus22:01

It's a good question. Let's sit in the world of hypotheticals for a moment and say it will occur — Sam Altman releases GPT-8 and it outperforms Harrison.rad.1. I think it's highly unlikely, but let's just say it does. The other thing about healthcare is what we call the innovation stack — my co-founder Dimitry loves this.

You need to build the models — the AI that's actually capable of subspecialist-radiologist performance. On top of that you need to build a product — one that fits into a clinician's workflow. It needs to receive images from the diagnostic system, it needs to display the result in a human-interpretable and safe way. All of this is regulated as a medical device — FDA, CE Mark in Europe, TGA in Australia, and many other jurisdictions. So you need to generate clinical evidence to prove it. You can't just publish a benchmark you ran internally — you need to publish in the Lancet, run a clinical trial to show the thing is safe and effective. And then you need to deploy, integrate in the real world, and do the change-management piece.

In the UK, a third of NHS England's chest x-rays run through our system. That's a journey we took, going with our partners at the NHS. And then, finally, a medical device eventually secures reimbursement. We recently received the first — sorry, not the first, but the only one right now — CMS reimbursement for our radiology AI in the US.

That's in the process of happening in every geography too, where the government essentially decides the technology is valuable enough that they'll just pay for it, rather than each hospital having to spend out of pocket. So if you think about what I've just described, each part of that stack is a layer of defensibility. The model is one huge part of it, and so far we've led the charge on that — but it's not the only part. So if I can see that model coming in the crystal ball, I'll just move further and further up the stack: into the productisation layer, the medical-device layer, and the reimbursement layer. Because then, in many ways, I can swap out a model any time. If they truly out-innovate us on AI — which our team will put up a good fight on — then we'll just take their model and deploy it in our clinical scenario. The goal of the company isn't to build the best AI. We're not here to build AGI. The goal is to automate healthcare and create infinite capacity. And whether we do that with our own model or someone else's doesn't make much difference.

But I still feel that even in that world, a fine-tuned version of GPT-4.8, with our own data and our own manually curated labels, would still perform better. So it's almost like: wherever that model gets to, we'll be better. The billions of dollars of investment coming into this ecosystem now really benefit us in a huge way.

Musty

So here's another thing — I haven't heard anyone discuss this, but I'm sure you've thought about it. This is inspired by Nick Bostrom's Superintelligence, or thoughts of AGI. My understanding is that if today you train your radiology foundation model on data labelled by humans — let's just say the best humans, through consensus — your performance caps at whatever the best radiologist could achieve today. But there's actually a layer above, which is supra-human performance: how can you be even better than a human? How can you spot things no human even knows exist in radiology imaging? Because I'm sure you'll agree there must be many such things — things we can't even interpret as humans that are in the imaging. We've seen evidence of that. Firstly, have you thought about that? And secondly, what would a world look like where you unlock superhuman performance? Would it be about using a different label for your ground truth? Would you have to rely on other indicators? How would you get to that world?

Aengus

Yeah, this is something we asked ourselves in 2019 when we started seriously doing data annotation: how do we outperform the underlying knowledge of the clinicians doing the labelling itself? And we're actually there today — we're outperforming radiologists now. I'd challenge any radiologist to sit down in a shootout with our model for lung cancer detection. They would not come out on top — statistically, by a wide margin. And that happens for two reasons.

Number one is consensus-based labelling: every case is labelled by multiple clinicians, so you get a higher level of sensitivity and specificity just from the wisdom of the crowd. More people looking at it — they miss it less, and they interpret it more finely.

But the second thing is what we call time-machine labelling. One of the neat things about our process is that we have historical data going back ten years. So when we show a chest x-ray to a labeller, we also give them the CT scan taken multiple years after that x-ray. If we show someone a CT brain, we also give them the MRI from the next day. When we ask them to label, we don't test them — we say, use whatever tricks you want. So they can look into the future. They can look at the future imaging and say, "Oh, that shadow there — would that have eventuated into cancer? That little speck — is that an artifact, or is it going to be a brain bleed the next day?" In a way, they're encoding the time machine into the label. This is something no human can do in clinical practice — there's no way to look into the future — but in labelling, you can. And because of that, our model isn't just predicting lung cancer on a chest x-ray. It's predicting whether the lung cancer will exist in this chest x-ray in the future. It's predicting the future. So that's some of the reason the system can be superhuman — beyond just what the radiologist is teaching it today.

Musty

That's incredible. That's incredible. Look, Aengus — if we put all humility aside — I'd love to hear some of the boss moves you've made in your journey. I really liked the story about building this 250-person labelling crew in Vietnam. That's awesome. But is there anything else that comes to mind? Any 4D-chess moves, smart moves — humility aside, I'd love to hear.

“You don't need a hundred AI engineers. You need a small crack team, and you give them a lot of compute — because that scales them a lot.”

Aengus

Aengus28:56

One of the things we do at Harrison that I'm pretty glad about is that we have our own compute cluster. Dimitry and I have this joke that whenever we raise a round of capital, we always buy a giant compute cluster. That's been true since our seed round. When we started the company, I bought a bunch of gaming parts from NVIDIA — the Republic of Gamers parts — to build a computer, to build AI. Since then, every time we raise capital, we always upgrade our compute capabilities. This was well before NVIDIA's stock went sky-high and it became pretty obvious that securing a lot of GPUs is a very key part of building the system.

But I'd attribute quite a bit of our success today to the fact that we haven't been reactive on the compute side. We've always secured — we have A100s in our server in Sydney, a large cluster of GPUs that we build and maintain ourselves. Which means we don't have a six-month wait time as OpenAI scoops up all the compute. That's been a pretty big enabler of some of the breakthroughs we have today. Harrison.rad.1 was all trained on our in-house cluster. We didn't need to go to the cloud for it. If anything, looking back, I think I under-invested — I should have bought a four-times-bigger cluster, because we'd have been able to do a lot more by now. And it's one of those compounding things.

I like to say a team of AI engineers needs to be a small crack team. You don't need a hundred AI engineers. It's not like traditional software, where a bigger team means better success. You need a small crack team of AI engineers, and you give them a lot of compute — because that scales them a lot.

Musty

And Monster Energy as well.

Aengus

Yeah, that's right, that's right. So that would be one of the things I did right. And I hope I get a chance to scale that again as part of the next chapter of the company.

Musty

One thing it appears you're pretty good at — if we go back to the IVF startup, and Harrison and the various products you've built — is convincing people to work with you. The reason I say that is, early on, building partnerships with large providers or useful partners. Now it's probably a little bit easier — you've got a bit of a track record — but I'm guessing at one point you were a medical student with not much to your name. There must have been a skill to going, "Hey, I've got nothing, but believe in me." I'd be very curious to hear your reflections on how you did it — some stories from that time. Because right now it's probably quite easy, but I'd imagine it was very, very hard back then.

Aengus32:45

One of the characteristics I identify with the most is that I'm a creator and a builder first — I always gravitate to that. That's the description I resonate with the most. I believe in building and showing that something can be done, rather than convincing or telling. Along the way, I've also been very fortunate to have great sponsors — people who make bets on me. If you read about me, there are different stories around the time in IVF, where there's a professor called Simon Cooke who came and gave a lecture on IVF and sponsored my first project.

But perhaps the advice I can give as a gift to the listeners is this: the agency behind what you do is very important. If you have an idea and you want to pursue it, just do it. Pick up the pen, or get out the keyboard, and start building a prototype — demonstrate that something can be done. It can be scrappy at first, but once you solve some part of the problem, it becomes contagious. Everyone wants to be part of building and solving great things. So when you see the early signs of that — and it's high activation energy — once you show it, then the momentum is very powerful.

And then people start to come in and want to solve other parts of the problem with you. But you need to activate that, by a big willingness to roll up your sleeves and get in and build the early solution first.

Musty

I hope you enjoyed that episode. And if you've been enjoying the podcast, then please do me two favours. First, if you could leave a review on Apple or Spotify, I'd be really grateful. And secondly, if there's anyone in your life who you think would enjoy this episode, please send it to them. Thanks for listening.