
In this episode of our Voices of Mathematics podcast we talk to James Fergusson, Professor of Theoretical Cosmology in the Department of Applied Mathematics and Theoretical Physics. James is also the Executive Director of the University of Cambridge's innovative new Master's programme in Data Intensive Science, and the Director of the Infosys-Cambridge AI Centre.
Fergusson's innovative collaboration with Infosys, a global consulting company and leader in next-generation digital services, is an outstanding example of how the University and industry can collaborate on cutting edge research with real impact and benefits for both sectors. "Out of our conversations with Infosys, we realised that a lot of the research challenges we have [in academia] are very similar to the challenges that large enterprises have," says Fergusson. "One of the really interesting things we are talking to companies about is how do you take these ideas from research and build them into robust, reliable systems that can be widely used."
As well as exploring the potential of what AI can do in industry, Fergusson is also equipping the next generation of researchers through the MPhil in Data Intensive Science. Rather than a more traditional Master's course which focuses on one specific subject, Fergusson says the MPhil is intended to be more like vocational training, giving students the full package of skills that researchers in this field will need.
We talked to James to learn more about the potential of AI to drive forward scientific discovery, how industry and academia can work together, and training the researchers of the future.
The podcast is hosted by Marianne Freiberger and Rachel Thomas, Editors of Plus, from the communications and outreach team at the Mathematics Faculty.
To find out more about topics mentioned in this podcast see:
- Our accompanying feature article Building bridges between academia and industry
- Find out more about the MPhil in Data Intensive Science
You can listen to the podcast using the player above, and you can listen and subscribe to our Voices of Mathematics podcast through Apple Podcasts, YouTube, Spotify and through most other podcast providers via Podbean. (The podcast is also listed without the transcript on the Maths Faculty website.) The full podcast transcript is available below. The transcript was created using AI to generate the text from the podcast recording, and was then sense-checked and edited for readability and accuracy.
[Musical interlude]
00:00:13 James Fergusson: [They’ll] sort of drive, the next round of, exciting breakthroughs. And we've sort of seen this, right, the Nobel Prize for Physics going to [Geoffrey] Hinton for the original machine learning thing. We've seen the protein folding and AlphaFold getting a Nobel prize. I think we'll see more of that, actually. These AI-powered tools will be driving breakthroughs in scientific areas.
[Musical interlude]
00:00:46 Marianne Freiberger: Hello and welcome to Voices of Mathematics, the podcast from the Mathematics Faculty at the University of Cambridge. I'm Marianne Freiberger and I'm from the Outreach and Engagement team here. The person you've just heard is James Fergusson and James was talking about the potential for artificial intelligence in scientific discovery. Now it might come as a bit of a surprise that James is Professor of Theoretical Cosmology in the Department of Applied Maths and Theoretical Physics, cosmology being quite an esoteric subject that investigates the nature and the history and the future of our universe. However, these days there is a wealth of observational information available to help cosmologists, which means they work at the interface of theory and data. And here artificial intelligence has proved extremely useful, both because it can analyse vast amounts of data and because it can help suggest theories that match the data. This relatively new aspect of his work has led James to embark on a number of fascinating projects.
In this podcast, James talks to my colleague Rachel Thomas about the Infosys-Cambridge Enterprise AI Center, which crosses the boundary between academia and industry. He also talks about a new MPhil degree in data intensive science, which he has helped to launch across the Department for Applied Maths and Theoretical Physics, the Department of Physics and the Institute of Astronomy here at the University of Cambridge. And finally, James also talks about an amazing tool called Dinario, which helps scientists to automate their research. And Dinario can even write papers - you'll find out more about this later on in this podcast. Here's James Ferguson talking to my colleague Rachel Thomas.
[Musical interlude]
00:02:41 Rachel Thomas: You're the director of something called the Infosys-Cambridge AI Centre, which is based in London and is a partnership with the company Infosys. What's the aim of that centre?
James Fergusson: So this came out of... we've known Infosys for quite a number of years via our PhD programme and some other things that we've been doing. And out of conversations with them, we realized that a lot of the research challenges that we have are very similar to the challenges that, you know, their customers have. Talking about those automation tools, they're obviously very generally applicable. And it comes from the fact that essentially data is growing exponentially, right? It's growing exponentially in the sciences with these enormous experiments like the Square Kilometre Array [SKA] that we put out, but it's also growing exponentially in every industry, in every sector, right? And so something that affects everyone is how do you get knowledge from data? How do you really learn something rather than just having these vast terabytes of data sitting there that you know is useful but you can't get the knowledge out of? And so this joint effort is really to think about how we are trying to solve our problems with these vast data sets that we have and the tools that we're developing and the ways that we're using machine learning and AI to do this, and then share that with people in industry and our large enterprises who have very similar problems with the large data that they have.
So, we do things, we have workshops, we meet the clients, we talk through our research, we try and write papers or articles about how they can be used. And then we have events as well where we can sort of give them a showcase of all the kind of state-of-the-art in this area.
00:04:14 Rachel Thomas: So you would take things like the sort of methodologies or techniques or tools that you've developed from your research, say in theoretical physics or other areas of research in the department, and you're sort of sharing those tools and exploring how they can be used in totally different settings that might have some kind of impact outside of the world of academia?
James Fergusson: Yeah, well, I think this is one of the really interesting things about the machine learning and AI world is that it's very subject agnostic. It doesn't really, you know, it doesn't matter where the data came from or what the area is. The tools you want to use tend to be fairly similar, right? You have to customize them a little bit, but on the base level, it is actually quite... the challenges tend to be quite similar. It's a way that you sort of see all of the world coming together a little bit. You know, biologists and chemists and material scientists and physicists and mathematicians are all thinking about quite similar things as, you know, large telecoms companies or manufacturing companies or that, because they want to do, they have these vast data sets and they want to do very similar things. They want to identify odd things in the data. They want to be able to classify them into groups. They want to be able to predict how they would, what would happen in the future. And then they want to also develop it, see if they can use it actually to develop a real understanding of the system that they're trying to study, whether that's their manufacturing supply chain or whether it's, you know, some particle physics experiment.
00:05:45 Rachel Thomas: And is there any of the examples from that partnership, from that work at the Infosys Cambridge AI Centre that you could tell us about just to give a sense of how this understanding can be shifted into different settings?
James Fergusson: That's quite difficult because when you talk to companies, everything becomes very secret very quickly. But I think one of the most interesting things, certainly the moment in the industry is working how these agentic systems can automate work. And I think that has a really general applicability. And one of the big problems industry tends to face is that they have a lot of legacy systems. They have a lot of data. The companies are old, they've been around for a while, they've used lots of different IT systems. The data will generally be fragmented, they'll be scattered across lots of places. And then they'll know they have a lot of it, but it's very hard for them to do something with it. And one of the really nice things about sort of these large language model powered systems is that they can do fuzzy cut and paste very easily, right? They can understand that the data in the systems are a bit different format or different way from the data on this other system, but it can work out, you know, because it's not complex, it's just annoying, right? It can work out that they're the same and it can bring them all together. So I think this is one of the big sort of, you know, undersold use cases.
You talk about the cutting edge, replacing people and all this kind of stuff, which isn't really there. But this really is, you can use this as a way of pulling all your data together into a place or making all of your data usable with these AI systems. And that's the thing that really then allows you to do all the really exciting stuff. I think that it's almost one of the biggest barriers to doing AI is solvable with AI.
00:07:35 Rachel Thomas: I think it's so interesting that idea of the two kind of cultures, the sort of academic culture, which is about sharing of ideas and it being public. It must be so interesting working with people who instead have to, for very valid reasons, have to operate in a sort of commercially sensitive, they can't share the work they're doing. Is that quite interesting operating across that boundary now?
James Fergusson: Yeah, I think it's really interesting. One of the things about being in the centre is I talk to... I now talk to a lot of people in business about what they're trying to do. And it's just fascinating to see the different kind of challenges they come up against, right? They often have regulations they have to worry about. They have to be very careful about how they do things. But a lot of it is just scale, right? So in science, you work with small teams of three or four people. You work on some data. The data never sues you if you get it wrong. Galaxies, if you misclassify a galaxy, it stays pretty relaxed about it. But in industry, that's not the case, right? They have to be really careful how they use the data and be really aware of biases and things that creep into it. But then they also have to do it - how do I do this when I have a million records I have to do or a million kind of things? Or I have to scale this across, thousands and thousands of people have to use these tools that I build - rather than academics who can deal with quite creaky and fragile systems because it's generally only used by two or three people who are experts in the field.
And so I think that's one of the really interesting things is talking to companies about how do you take these ideas from research and how do you actually build them into robust reliable systems that are used, they're able to be used widely. And it's sort of, it's becoming more of an interesting thing for academia because as we build these tools, they are generally applicable. So if we build something to do cosmology, it may well be really good at doing chemistry or biology or social sciences data analysis. And so then we have to think about those same challenges. How do we make it robust and usable for people who are less expert or maybe less familiar with these tools to still get really good value out of them?
00:09:35 Rachel Thomas: Because I was thinking, it's more straightforward to see how beneficial this is to the sort of commercial partners of this. That's an obvious thing because they get to access cutting edge research by the department and people like you. But it's interesting you said that. So, is it interesting for you as a researcher to be working with these people or working in this way?
James Fergusson: I think it is. It's interesting because the challenges are quite similar. They do have quite interesting problems. But also thinking about the deployment side of it, working with someone like Infosys, who's an expert in this, they have huge teams of software developers who are really, really, really good at taking ideas and deploying them in a robust and scalable way. It's really interesting for us to learn about like, how do they think about building these things? And how do they think about building these things at scale? It is something we don't normally think about, but we're starting to have to think about.
00:10:35 Rachel Thomas: So as well as sort of practical implement... as well as practical questions that has academically interesting questions because of this idea of transferability of research.
James Fergusson: I think so. I mean, the one thing about the AI space is that it's one of the few spaces where you can, you can quite easily be in a company doing, publishing lots of papers, exactly the same as a researcher, right? There's very little boundary between being an AI researcher in academia and being one in industry in terms of, whereas in any other field, you would, say in material science research, there'll be companies that do good material science research, but the best research will probably, you know, be in academic environments. And that's true of most scientific areas. But I think AI is the one area where, you know, it's not. Industry, you know, there's lots of great stuff happening in academia, but industry is also doing really, really great stuff. And they have vast resources to do the really big scale things that academia would normally struggle with. And so they're very interesting to talk to because they actually are doing cutting edge things in industry as well.
[Musical interlude]
00:12:00 Rachel Thomas: So you mentioned something called Denario. Is that an sort of AI tool or system that you've been involved in developing?
James Fergusson: We've been building systems like Denario, which is just coming out, which actually does full end-to-end data analysis in a fully automated way, right? So you put the data in and it will generate research ideas, it will code up all the code, it will run the code, it will get the answers, it will write the papers, it will do all the plots. And you can use this essentially as a convenience tool for saying, I want to ask these questions of the data. And rather than going through the boring grunt work of writing up your Python notebooks and then getting all the data and doing all your plots, which can be quite time consuming, we can now automate a lot of that. And so we're using it in a similar way to [how] a lot of people in industry are using it, which is really as a time saver, as a way of automating sort of repetitive and reasonably straightforward tasks that you do. And even key researchers spend a lot of their time doing these.
On the theory side it is also useful. So we're using it on that one. You can use it to explore solution spaces of equations efficiently. It's something we've been working on for a long time in the group.
Because we've sort of seen the power of these tools, you know, we've seen it for writing code, right? You know, we now know that AI essentially is better writing the code than almost anyone in the world. It wins all the coding challenges. And a lot of data analysis is essentially just about writing code. It's trying to get your mathematical understanding and turning it into analysis. And so we've been working for a long time on how do we use this in more and more advanced ways.
And this is a multi-agent system that actually allows about 40 different agents, which are fancy hats that you give to large language models to tell them to do specific tasks. And you sort of put them together like a team of not very bright, but very efficient people. And you can then build quite complicated pipelines where they can do, quite sophisticated analysis. We've used them... they're great for engineering challenges. If I want to build an emulator for a system, something like that, it can go out and code up all possible options, train them, test them, and then come back and tell you, know, this is the best one for the system in about, you know, 10 minutes.
Rachel Thomas: So it's like an assistant rather than replacing the researcher.
James Fergusson: Yes, I mean, the thing, they can't, they're not intelligent, right? This is the thing that, we talk about artificial intelligence, but they're not, right? They're not intelligent. They can't think, they can't really do any kind of judgment or analysis or come up with any new brilliant ideas. What they can do is do things they've already seen quickly and efficiently. So if they've seen lots and lots of code, so they can write code very quickly and efficiently. They've seen how to do a lot of statistical analysis, so they can implement that quickly and efficiently. But they're not doing anything that we couldn't do. They just do it faster.
00:14:54 Rachel Thomas: And can you give us like a simple or a small sort of test example maybe that you've used Denario on, like an example of some research you got it to do?
James Fergusson: Well, yeah, so we use it for these like engineering type tasks. We have recently tried to sort of test the limit of it and we gave it a whole bunch of data sets from across all sorts of parts of science and it coded 80 papers automatically, just try and do them all. And then we tried to send those papers to people to review them. And as you'd expect, a reasonable chunk of them were not very interesting. It wasn't doing anything particularly original, but there were also, some really good papers that came out of it. Because it can essentially, you can just, if you give it data and say, well, what are all the ways you can analyse it? Whereas a human could only try two or three ways, it can try all of them and then it can compare the results and then tell you what's best. So that exploring is where it's really, really useful, I think. Exploring all the possible methods, all the different machine learning architectures or tools you could use to analyse this data and just implement all of them and then review which one performs the best.
[Musical interlude]
00:16:15 Rachel Thomas: And you're now very involved in amplifying and exploring the use of AI in your area of research, but also in equipping the next generation of students coming through the university with those sorts of skills. One thing I know you've been involved in is this new MPhil in data intensive science and you're involved in, you were talking about working with PhD students, some of whom have been involved in Infosys as well. Why do you think it's important? What are you hoping to give to those new students training now as researchers? What do you hope they take away from the current state of AI and where do you think they might take it?
James Fergusson: Yeah, so these programs we developed essentially because of this explosion in data that we have. And that people were coming to us with traditional training in physics and maths. And they'd be brilliant at doing calculations, but then essentially we're giving them research, which becomes almost an engineering type task, right? If you want to analyse a 10 terabyte data set, then you actually do need to know a certain amount of computing. You need to be able to program pretty well, and you need to understand how you can use these very scalable machine learning methods to work on this data. And so we developed this MPhil out of our PhD programme as the set of skills we thought, you know, someone who wanted to work with data in the physical sciences would need. And so I almost see the MPhil as like a vocational training, whereas it's very, it's quite unusual in that it's quite different to a traditional Masters that we'd offer, which are normally subject specific. So we'll give you a Masters in statistics and you learn all about statistics. And in this one, you learn a bit of statistics, you're doing a bit of machine learning, you're doing software development, you learn high performance computing, and then we also do applications to specific research areas to try and give you that breadth of how do you actually implement this in the real world. And so my great hope is that people who come to us and come to on this Masters, they have their great training from their undergrad, and then they can use this sort of vocational skills training of how do you work with data and how do you use these really cutting-edge tools to sort of drive, the next round of, exciting breakthroughs.
And we've sort of seen this, right, the Nobel Prize for Physics going to [Geoffrey] Hinton for the original machine learning thing. We see the protein folding and AlphaFold getting a Nobel prize. I think we'll see more of that, actually, these AI-powered tools will be driving breakthroughs in scientific areas. And people, you know, mathematicians and physicists are always very well placed for this because they tend to find understanding the machine learning tools very easy because the mathematics of them is generally linear algebra. It's reasonably straightforward. But they tend to be the ones who can really drive forward these tools. And they're the ones that industry also really want to hire. And so I hope that they'll go out and build the next generation of scientific tools that should hopefully then be used across a wide range of areas.
00:19:23 Rachel Thomas: And can you tell me about the MPhil, what kind of people are coming into that program and what sort of, what's their experience like doing it?
James Fergusson: So one of the things about this program being a little bit different to the traditional one in that it's a broad range of skills is it attracts a very broad range of people, much more so than we traditionally get in some of our programs. So we tend to, we have 46% women on the program. We have some of the largest age ranges. We had a 61 year old did it last year. We have a lot of people coming back from industry. We have a lot of people coming from different fields as well. So we're mostly physicists and mathematicians, but there's also other mathematical scientists and software engineers and engineers come. And so we see one of the really nice things about the MPhil is it's quite like a melting pot. It's quite a rich cohort of interesting individuals from a lot of different areas. And we see this when the students go through it, they learn an enormous amount from each other, from being exposed. You know, there may be a physicist there who's talking to a professional software engineer who's come back from industry who's talking to a mathematician who's come from a quite pure area to maybe a machine learning engineer who's been using these tools but wants to get better foundations. And so we see a huge amount of sort of inter-cohort learning as well, which I think is one of the really nice things about this program being different from other programs which tend to have very homogeneous cohorts. And as part of that, we've done a lot of work to try and encourage this. So we have lots of, we have team building days. We have a psychologist who comes in and works with them to talk about how to get a mindset for success. And we also have a lot of training on presentations and writing and all the other kind of skills that you need to be a successful scientist to try and bring that group together and really try and give them all of the skills that they need, the academic ones, but also the soft skills that they need to be successful.
Rachel Thomas: It sounds like such an amazing programme. I'm not surprised you have a lot of people wanting to come on it.
James Fergusson: It is very popular, but there's always space for anyone interesting.
Marianne Freiberger: That was Rachel Thomas talking to James Ferguson here at the Centre for Mathematical Sciences at the University of Cambridge. If you have enjoyed listening, then please consider recommending us to a friend or rate and review us wherever you've been listening. I'm Marianne Freiberger. Thanks for listening and bye bye.