Mustafa Ghafouri

What’s in this episode

An impressive AI demo can make production feel like a formality. But the leap from a controlled sandbox to a system that delivers real value is where most AI initiatives stall.

In this episode, we speak with Mustafa Ghafouri - Head of Data Science at Pharmacy2U - about why so few AI prototypes make it to production, and what the ones that succeed get right.

Mustafa left his medical degree for data science over a decade ago. Since then, he's consistently shipped models into production across healthcare and tech, including as a solution architect at Databricks and on projects with NHS Digital. He explains why operationalising AI is the most underestimated part of any project, which questions he asks before he believes a demo, and why he builds the solution before the model. We also look at what made an NHS content-moderation project work, what has held back ambitious ideas like AI scribes and automated triage, and why simple use cases done well so often beat grand plans.

If you're about to launch your first major AI initiative, or you're stuck somewhere between promising prototype and real-world impact, this is a conversation for you.

Meet the guest

Background pattern

Mustafa is an accomplished data science leader at Pharmacy2U, with a solid track record of taking AI from experimentation into production. Having trained in medicine before moving into data science, Mustafa combines deep domain knowledge with hardcore technical expertise, bridging the gap between strategy and technical delivery, and ensuring engineering gets a seat at the table.

Mustafa Ghafouri

Transcript

Yasin: Welcome to Ctrl. Alt. Deliver. Today's guest has spent the last decade sitting at some of the most important intersections in healthcare, data and AI. He's worked across the NHS, Hitachi, Capgemini and Databricks, and now serves as Head of Data Science at Pharmacy2U.

What's interesting about Mustafa's perspective is that he's seen AI from almost every angle: as a clinician, as a regulator, as a consultant, as a data leader, and now as someone responsible for deploying AI into products and services that affect millions of patients.

And that's exactly what we're going to explore today, because there's a huge difference between an AI demo that impresses people in a boardroom and an AI system that survives regulation, integrates into workflows, gets adopted by users, and actually delivers value at scale.

Mustafa, welcome to Ctrl. Alt. Deliver.

Mustafa: My pleasure. Thank you for having me.

Yasin: So, Mustafa, you've had a very unique career. You started off as a doctor, and today you're leading data science teams and AI initiatives. Tell us about that journey, because that is truly a unique one. And what advantage do you think it has given you over people like myself who've only worked on the tech side?

Mustafa: That is a great question, and one that I get asked quite a lot. Despite that, I never have a pre-prepared answer.

Like a lot of people, when I left college, I wanted to become a doctor. However, during my university degree, I realised it wasn't as enjoyable as I would have liked. So I was looking for alternative careers, but still in healthcare, because I was still very passionate about healthcare. At the time there was an article – I can't remember if it was Forbes or somewhere else – titled "Why You Should Get Your Kids to Be Data Scientists and Not Doctors." For me that was very relevant, because I was nearly qualified as a doctor at that point. I thought, what is this thing, data science?

I'd always had a little bit of resentment that I wasn't able to study maths, statistics and computer science at university, which I think I was more naturally affiliated with. So I thought, okay, I'll look into this. When I saw the master's – it was the first year they were doing it at UCL – I thought, this is a great bargain, because you're getting to study the maths and the computer science, learning to code at the same time. I didn't think I'd change careers into data science as such, but I knew it would open doors. However, as soon as I did my master's, lots of opportunities came flying in, and the rest is history. I think that was 10 or 11 years ago now, and I've been in data science ever since.

The change in landscape over the last 10 years has been immense, from [unclear] nervousness to what we have now. I remember my parents thinking I was crazy. They were like, why are you leaving this stable career in medicine to do this thing that no one's heard of? And obviously, 10 years later, that's all flipped. Back then, on my data science master's course, there were 30 places and I think only 13 people applied. There were nine internships and only three people applied for them. Whereas now – I checked recently – it's more competitive to apply for a master's in data science at UCL, I believe, than it is to apply for an undergrad in medicine.

Yasin: Oh wow.

Mustafa: Yeah, and 10 years ago, that just wouldn't have been believable.

Yasin: Given that you were a doctor and then got into data science, do you think your medical background helped you quite a bit? Because you need to know what the data is that you're manipulating, right? So I'm sure that gave you a very good basis for jumping into the technology side as well.

Mustafa: Yes and no. They often say that defining what a data scientist is is actually quite hard to this day – who counts as a data scientist. People often say you need to be better at maths and stats than a computer scientist, and better at computer science than a mathematician. But the third pillar of data science is always domain knowledge. The reason for that is, unlike maybe a data engineer or a pure backend engineer, you have to be much closer to the business problems that you're solving.

So without a doubt, understanding the business problems in healthcare – whether that's demand and capacity for hospitals, genetic profiling, disease risk profiling, or even infection control in hospitals – was a lot more intuitive and easier for me, and I was much more naturally able to fit into those business conversations.

I would add, however, that I had a massive learning curve in terms of business acumen, because in the UK you get trained under the NHS, and I don't think I quite got the business training. I had to learn that a lot more on the job. But understanding the business problems and where the value was was a lot more intuitive.

Yasin: Yeah, definitely. And that brings me to the next question. You're Head of Data Science at Pharmacy2U – we all know it's a very popular online pharmacy, one of the leading ones. When we last chatted, you talked about going deeper than distribution: helping patients manage their medication and improving adherence to it. How much of a role does AI play in that vision?

Mustafa: It's a great question. When I joined Pharmacy2U, I was the only data scientist. It was a new initiative that the company was embarking on, and I think we've done really well. I've been very fortunate to be at Pharmacy2U in the sense that they've given me the opportunity to really deep dive into some of the business problems and the ways in which we can help patients, and to begin this digital and AI transformation journey that we've gone on together. And it is a journey. Obviously it's not just me – there are other people, and we all have to move together and learn progressively. It's not the sort of thing where you switch a light on overnight and you're suddenly flying out of the gates. There are a lot of lessons on the ground.

In terms of the actual use cases, I've been really happy that we've been able to cover a lot of areas. You're correct that medication management is an area we're actively looking at, simply because every month we get a lot of prescription data. That's a lot of data to understand every month, so being able to help patients in that respect has required AI.

Another area is anything where there's unstructured text – areas where there are huge amounts of it, whether that's feedback we get from different sources, et cetera. Because it's such a large company, we do get a lot of feedback from customers, and being able to analyse that at scale and understand what people are saying is another area where, without AI, it simply wouldn't be possible at that scale.

Yasin: I think we can all see that the potential of AI in healthcare and other industries is significant. You've spent 10 years in the weeds and the trenches. What's the biggest thing you would say we're underestimating about AI?

Mustafa: Great question. By far, for me, the biggest thing is operationalising AI. There's a lot of focus on building the models and getting your accuracy or performance metrics up to a certain level, with supervised learning models and so on. However, actually getting value out of a productionised system is very different to building a model, or even fine-tuning a model, or writing a system prompt for an LLM.

A lot of the time in companies, what I see is people will build a model and go, okay, it's got this accuracy, we can do this with this accuracy – in an isolated sandbox environment. And the lesson I see companies having to relearn every time is that going from that isolated environment into a productionised system that's actually adding value is the hardest part. It's the part that, when you look at project management plans, can often be the most underestimated. Recently we've had a lot more honest assessment of that – machine learning operations is now a field of its own, and machine learning engineers who are responsible for this are a very prominent role – but overall [unclear], I think there's a lot of progress to be made on that.

People often talk about the gap between big tech companies such as Google, Facebook, et cetera, in implementing AI, and everyday companies doing business in various domains, whether it's healthcare or otherwise. The big difference I see is that the big tech companies know how to operationalise AI quite well. They make it a key focus, and the engineering is very high quality. Importantly, there's a real emphasis on good quality engineering that you simply don't get in healthcare companies.

I think it's misunderstood a little bit. If you go to your average NHS trust or integrated care board now, a lot of the people managing the tech projects are not technical, and I do find that a real concern, if I'm honest. There just isn't this appreciation that business decisions and management require a deep technical understanding – especially in AI. To productionise the models, people think, okay, the model's deployed, and they don't realise, for example, that there's drift in the models that might deteriorate over time. Or the way the model's trained might not be appropriate for actually producing results. Or maybe you're not even measuring your model correctly in terms of the value you're having.

There are all these issues in operationalising models, and a lot of them are quite scientific and statistical in nature, which might not come out on a project plan. A project plan might say, "Okay, we've deployed the model," and no one's checked that it's accurate. No one's checked that you're actually improving what you said you'd improve. There's almost this blind trust that happens. And then when the numbers don't go up – if you're not helping more patients, or if your finances don't look better – either it goes unnoticed, or people question it and struggle to get to the bottom of why things haven't improved.

Yasin: One thing you pointed out, Mustafa, was a lot of CTOs these days being non-technical. We run into quite a few of them – that's probably a topic for another discussion. But now the stakes are getting so high that engineering excellence and an appreciation for technology at a much deeper level is key and critical for these systems. A lot of them do come into production, but you don't have the iteration to improve, or the appreciation of how to improve and why.

Again, that leads to the discussion that's getting into the headlines as well: there are a lot of AI demos out there, and very few of them get to production. There are a lot of people dabbling and experimenting with AI, but there seems to be a huge gap between the demo that everyone gets excited about and what actually makes it to production. Why do you think that gap exists, and why is it so difficult to close?

Mustafa: One thing I'm very proud of in my career – in my 10 or 11 years of being a data scientist, and I've also been a technical solution architect at Databricks, and in other roles such as data architect and data engineer – is that I've got a lot of stuff to production. It's something I really pride myself on. Depending on the details, I think there's only one contract in the last 10 years where I haven't gotten something into production.

That's very rare for a data scientist, believe it or not. I talk to a lot of other data scientists at companies, and none of their models are deployed, or they'll spend time building a model but it's not actually producing value. Echoing what I was saying earlier, a large part of that is understanding how to productionise models. There's a real tactile skill there that is not appreciated at all.

A lot of the time when I go into a project, particularly if I've got a team, that's all I'm focusing on: how to productionise the models. In the same way that if you build a car, you want good quality engineering – when you're planning how to build a car, you want to understand the different parts, the engine, the exhaust and how it works – in my mind, data science works in a very similar way. When I look at a project, I really look at what good quality engineering looks like, whether that's data engineering or otherwise. Otherwise you're just going to have problems and you're never going to get into production, and you need to get ahead of these.

A good example is that I will often try and build the solution before the model. What I mean by that is, I'll build at least part of the environment in which a model is going to sit before actually starting the experimentation phase of building the data science model. The reason for that is you then start to understand the context within which the model has to operate. So you can start understanding efficiency problems, latency problems, problems in what data is available when. A lot of people accidentally train models on data that wouldn't be available when you need to call the model. It's a very common problem.

Just lastly: you can have non-technical CTOs, and I think there's a really important role for them. But at the same time, it's not a surprise that a lot of the biggest and most successful companies in America have technical CEOs – not even CTOs, CEOs. I'm not saying everyone should be technical, but whether you're looking at Microsoft, Meta and so on, it's really not shocking that a lot of the companies that successfully implement AI have technical CEOs, because they just understand what actions you need to take, and the landscape and the context, so much better.

I don't even think it's common knowledge that operationalising AI is the hard part. If you go to the average CEO of a FTSE 250 or FTSE 100 company, I'm not sure they would know that. In their mind it's, "I need to make business decisions, and the techie people can deal with the techie stuff and just tell me what I need to do." That is true, but if you're going to make strategic decisions, there is a really large technical element to them.

I should also be clear that a lot of people don't agree with this view. It is very common to hear, particularly from non-technical people, that the tech's the easy part and the business stuff's the hard part – the people side. Don't get me wrong, that is difficult. But you would never say that if you were building a car. And this is what shocks me: if you're building a car or doing any other engineering, such as building laptops, you would never say, "You don't need to understand the technical stuff, they can just do that in the corner, and I just need to know the outcomes." But in a strange way, that's what happens with AI a lot of the time. And this isn't one or two companies – it's been the experience for the vast majority of my career.

So a large part of my job is always taking people on a journey to understand that it's not just something the techies do in the corner. It's a major part of the business discussions that need to happen, and it needs a seat at the table.

Yasin: In our experience, what we've noticed is that all of these AI demos and prototypes are done in such a controlled environment. There's a sample dataset, and it represents a portion of the actual data. When you need to deploy it, there are 10 different data sources, so many different integrations, and the data isn't normalised. So a lot of the time, what we've seen is that taking it to production just stops there. Cleansing and normalising the data is a much bigger job than the AI part itself. And then they're like, "We're going to spend eight months just doing that, and then we'll get to AI" –

Mustafa: What's even worse is that sometimes people aren't even aware what the issue is. So projects get delayed indefinitely, because you highlight a technical issue and it just gets ignored. Latency is a great example. The efficiency of models, calling the models and the backlogs that could cause – all these sorts of things just get ignored a little bit. And it's understandable. I think the difference is that if you're building a car, people can see the parts. When you're building AI models, it has the same principles and the same parts, but because people can't see them, I think that's why they get ignored.

Yasin: It's very interesting – all the normalisation of the data and so on.

Mustafa: Yeah.

Yasin: So, Mustafa, when you see an impressive AI demo, what's actually going through your mind? And what questions are you asking yourself before you believe it could actually work in the real world?

Mustafa: Great question. When I see a demo, I'm thinking: okay, what source data did they use? What process is that going through? The biggest one is probably: is this applicable in a wide variety of dynamic situations? That's probably the biggest thing, because the reason demos work is that they work in a very narrow set of conditions and contexts, and they're very controlled for the purpose of the demo.

A great example: you could literally write a script that prints out a made-up conversation. I might do a demo of a chatbot for, say, health or social care, and every time I press enter it prints out the next line, and it looks like an amazing demo. But the problem is it only works in that very specific context. In real life, no one's going to have that exact conversation. The whole advantage of tech is that it's dynamic across a lot of different contexts and situations.

So what I'm thinking is: what situations and contexts haven't they presented, and would it still work then? If it's a chatbot, I might think, okay, it's working when they ask this question – what if they ask this? What if they tried to break the chatbot? What if they swore at it? What if they tried to get some personal information out of it? So I'm really going through the different contexts.

Additionally, I'm thinking about the engineering. There's a limited amount you can get from a demo unless you ask someone technical, but I'm thinking: have they thought about the data architecture? Have they thought about the data pipelines properly? Have they thought about integration? Have they thought about security? And so on. So I'm going through almost a checklist in my brain of all the different contexts and situations they haven't presented, but also all the system engineering principles as well.

Yasin: System engineering – I think that's the key word here, and it's what a lot of people miss when they're doing these AI prototypes and demos. The AI might work in a controlled environment, but it's all the system engineering around it that needs to happen for it to actually be a success. So when you've seen projects make that journey successfully from prototype to production, what were the key stages they had to work through?

Mustafa: I'll be honest, I don't think it's a case of going through specific stages. I think one of the most successful principles for productionising models is small agile teams. One of my most successful projects, in my opinion, was when I worked at NHS Digital, again as a contractor. I helped them with a project on moderating content on their website. If people were posting reviews that were abusive or had confidential information, we wanted to identify that so it wouldn't be posted. The reason I really like that project is that it was a really nice, close, intimate team. We worked in proper agile, and that was lovely.

If we think about the different stages, I think that's useful, but I don't think it's the absolute key thing. One is, like we said, system engineering from the beginning. Thinking about that NHS Digital project, I co-led a small team, and we had different data scientists working on the models. But my role, apart from leading strategic conversations, was largely the system development. Because I'd worked as an architect previously myself, I worked with the solution architect and the enterprise architect to design the solution. So I was managing both at the same time, and I think that pattern works really well. I was essentially acting as an ML engineer – almost like an AI architect – as well as leading the data scientists doing the work.

What that meant is that, in parallel, we designed the context within which the models would run, and that allowed us to give the data scientists everything they needed to know to actually build and design the models. There was –

Yasin: So, Mustafa, was that data for the whole NHS, or for a particular trust?

Mustafa: No, this was for the NHS website.

Yasin: Oh wow, so we're talking about large data. All GPs.

Mustafa: It wasn't really about the large data. I mean, it was a large dataset, but it was just the NHS website. When you go on the NHS website, you can leave a comment or a review, so it was for those sorts of situations. If anything, we didn't have enough data, because we didn't have enough labelled data.

So a great example of system engineering: we'd got so far with the system engineering we'd done initially, but then we were thinking about how to get more data, and we couldn't just ask people to label lots of data. It's a great example of how strategic thinking and technical knowledge have to play together. In order to plan it all strategically, you have to understand what type of data you need, the fact that it needs to be labelled, how it needs to be labelled, the quality, what performance metrics you need as a data scientist, et cetera.

That's where tools like GPT-Neo came in – it's not very popular now, but it was similar to GPT-3 at the time. They helped us improve our engineering and the data we had – for example, synthesising data we could use to train models. So all this strategic thinking became a huge part of my role, and on the whole I would say we were very, very successful. We integrated with the NHS website quite cleanly. We did it using a parallel processing system, with multi-threaded API calls to the models. We were considering aspects such as latency – when the models could reply and how long it would take – and the quality of the replies, what was confident, what was good, and how things changed over time

Those initial conversations I had with the product owner, who was the business lead and was absolutely brilliant, and with the technical architect for the NHS website at the time – they were all very positive. There were no barriers. When I said, "Okay, I need to talk to the technical architect," there was no hierarchy. No one said, "Well, you're just a data scientist, you can't talk to the technical architect."

The thing I think really made that project successful is that the product owner himself was very respectful of the tech experts. That was the absolute key thing for me. There was this real appreciation.

Yasin: Yeah, appreciation.

Mustafa: Which you don't always get. Some people think the tech people are just there to build something, and they don't really want to interact with them. He was very engaged with what we were saying, and really engaged with the caveats, because there were a lot of quite detailed business decisions we had to make about what was moderated and how it was moderated. So it was a real intersection of tech, particularly data science, and the business problem.

As an example, if you want to prevent someone using descriptive words against someone – someone might say "the blonde nurse was x, y and z" – you want to prevent that. But to what degree do you prevent it? Because they might be saying, "I'm a blonde person and my hair was dyed black," or something, in the comment.

Yasin: Right.

Mustafa: I mean, terrible example, but you get what I mean. So there's a real interaction between the product owner and the business problem and the tech problem, and that requires a constant two-way dialogue, which just happened. And it doesn't always happen. Sometimes everyone's in their own little world trying to fix their own problems. Whereas we really worked as a team on that problem, and I think the results speak for themselves.

Yasin: Good. So, last few questions. Without naming names, what's the best AI idea you've seen that never actually made it to production, and what killed it?

Mustafa: That is a very tough question. I'm going to say a few things that probably contradict some of what I said previously a little bit. I'll talk about two use cases, and they died for different reasons – or rather, they're ongoing. These are actually very successful use cases, but they're probably not as prominent as they should be.

One is AI scribes. I don't know if you know about AI scribes in healthcare?

Yasin: Yeah.

Mustafa: So, documenting the interaction between the doctor and the patient, and then automating things off the back of that. I'm not up to date with the latest in this area, but that might mean writing letters, doing prescription recommendations, doing all the admin around it to make the doctor's life easier. When I last checked, which was a while ago, that wasn't doing very well, for different reasons.

I think the technology aspect was interesting, and people had integrated it quite well, actually. But I think there were challenges around, one, regulation, because it wasn't very clear what was happening – that's an interesting example where the regulator was playing catch-up a little bit. But more importantly, the value add hadn't been properly discussed, at least a year ago. I think it's better now. What I mean by that is, the tech wasn't at the stage where you could be very confident in what was happening, and for some of these actions you needed to be incredibly confident. If the doctor is having to go through everything the AI has done and double-check it all, and the use case is all about saving the doctor's time, you're not getting value. So at least a year ago – it might be different now – that was what killed that use case.

It's a use case that so many different companies were doing in so many different countries, whether it's France, America or the UK. And I think the thing that would have saved it – and maybe has – is a real focus on: okay, what can you do accurately, and where can you get value? Is it on doing a certain type of letter? One of the best use cases I've seen is from a company some friends of mine run, called Health Tech One. The first thing they did was just automate GP registration – when you sign up to a new GP. That was their entire thing for a while. It's a complex enough task that it takes admin time, but the value add was simple, and they did it cheaper than the GP practices could do it themselves. So they made it a no-brainer for the GP, because the value add was so clear. Getting to the heart of that value add can be difficult.

The second one, which is very different, is diagnosis. I'm thinking of Babylon Health here – I'm sure most people will know the Babylon Health story; if not, you can Google it – and in particular triage systems. Interestingly, I think the NHS is having another go at this themselves; I think they're getting a consultancy to do it. It's been the miracle dream of health tech to do triage systems. It makes sense, right? If you've got some symptoms and you can explain them, and you can triage automatically, it'll save so much capacity in the system, which can be used elsewhere. That's the value of it.

That one's an interesting one. I think it struggled, and is struggling, for multiple reasons. One is that people are too ambitious with it. You need to be safe with this sort of stuff. And the reality is that a doctor who has a conversation with you, firstly, feels safer, and can get more out of you than an AI system might be able to. A human doctor just has that intuition. They can see your face. They can see your body language. They can see if you're nervous or anxious. I think that's one of the reasons it's really struggled. The other is that the tech maybe just isn't there yet to do the full, ambitious plan. So in that example, I think a leaner, safer triage system is probably where things are going to go. I don't know what the NHS app is going to do, but I believe they're investing quite a lot of money in this area.

What are your thoughts on those two use cases?

Yasin: My thought is that technology is moving at a very fast pace. We're doing this podcast today, and we might come back to it six to eight weeks later and a lot of it might already be outdated. Things are changing very fast. The other thing you touched on is that the need for a human in the loop is still there, and there's a very high likelihood that it's going to stay.

And there's that whole discussion – again, another discussion – about whether AI is going to take our jobs. A very dense discussion. But in my opinion, jobs are evolving. They have evolved, and they will evolve in the coming months as well. They won't disappear; it's just the way you do them that's going to evolve. Let's put it that way. But again, these are all very dense discussions. Maybe we'll catch up again on one of these topics and talk about it in detail.

That brings me to the last question. I've taken a lot of your time already. If you're advising somebody who's embarking on their first major AI initiative today, what's the one thing you'd tell them to focus on to make sure it actually gets to production?

Mustafa: I like simple use cases. I like simple things done well, and that would be my number one piece of advice.

When we met, I'd just launched Hospital Waits, if you remember. The day we met, it was exploding. I'd released it a few days before, and I think it got 12,000 visits in a week. I've reflected on why that is. It's not explosive growth in the sense of what Facebook or a company like that got, but for me it was quite impressive. I just shared the link on LinkedIn, and all of a sudden 12,000 people in a week were using this website that I'd built in a few hours using AI coding tools.

I think part of the reason is that Hospital Waits is an incredibly simple tool. For those who don't know: if you need an operation, you go on the website and put your postcode in, and it tells you the wait times. If you see a hospital with short wait times, you can put in your information and it'll write you a letter you can print out and give to your GP, using your right to choose, so you can be referred to the hospital with the lower wait times. I know it's not quite true, but I aimed for it to be so simple that your grandma could use it.

The projects I see fail are the ones where there have to be so many interactions and touchpoints for it to succeed that I think it's just too complicated, especially given the complexity of the business world, and especially if you need more stakeholders. So for me, simplicity and effectiveness are absolutely key, and that's the challenge. It's the use case. Do you have a hundred stakeholders who need to align perfectly for this to succeed, or can you build it in a simpler way, where you only have a small team that needs to work together to build it successfully? And is the use case itself something where, if you tell your business stakeholder in a sentence, they can understand the value?

Yasin: Very key point, Mustafa. Keep it simple. Go for a very simple use case. Do it right, do it well. And more importantly, do it with a team that can be fed with two pizzas, like Mr Bezos says.

Mustafa: Yeah, that's small enough to do it. He's spot on.

Yasin: And the thing is, don't boil the ocean. You can boil the ocean later. Just do it one step at a time.

Mustafa: Yeah. Just because you've done something simple doesn't mean that's the end goal and nothing's going to get better afterwards.

Yasin: Good. Well, Mustafa, thank you very much for your time. It was very informative. Audience, it was a pleasure hosting Mustafa, and hopefully you all learned how to take AI from prototype to production. Thank you, Mustafa.

Mustafa: My pleasure. Thank you very much.

Ctrl Alt Deliver

What would you like to hear next?

Have ideas for new episode topics or guests? Tell us your suggestions or feedback.


      Build for what's next.

      Tell us about your systems, challenges and ambitions.


      By submitting this form, you agree to GoodCore Software Privacy Policy

      20+

      years delivering exceptional software

      100+

      critical systems delivered


      Check Mark
      NDA Included

      Strict adherence to confidentiality

      Check Mark
      IP rights secured

      Intellectual Property belongs to you