skip to content

Features: Faculty Insights

 
Voices of Mathematics: Richard Samworth

In this episode of our Voices of Mathematics podcast we talk to Richard Samworth, Professor of Statistical Science and Director of the Statistical Laboratory in the Department of Pure Mathematics and Mathematical Statistics (DPMMS). 

 

Richard Samworth's outstanding contributions to statistics have been recognised with numerous honours and awards. In 2025, he won two prestigious prizes in the space of 24 hours – the David Cox Medal for Statistics and the Guy Medal in Silver.

He has also been invited to speak at the International Congress of Mathematicians (ICM) in July 2026. Held every four years, the ICM features the world's leaders in the field and celebrates the diversity of today’s mathematics. Only mathematicians whose work is of the highest international standard are invited to speak at the week-long event, which usually draws thousands of participants working in all areas of maths. 

We talked to Richard to learn more about some of the most exciting developments in his field of statistics, the work he will be speaking about at the ICM, and what he values about being part of the mathematical community in DPMMS.

 

The podcast is hosted by Marianne Freiberger and Rachel Thomas, Editors of Plus, from the communications and outreach team at the Mathematics Faculty.

To find out more about topics mentioned in this podcast see:

 

You can listen to the podcast using the player above, and you can listen and subscribe to our Voices of Mathematics podcast through Apple Podcasts, YouTube, Spotify and through most other podcast providers via Podbean. (The podcast is also listed without the transcript on the Maths Faculty website.) The full podcast transcript is available below. The transcript was created using AI to generate the text from the podcast recording, and was then sense-checked and edited for readability and accuracy. 


[Musical interlude] 

00:00:14 Richard Samworth: DPMMS is a fantastic research environment. I think the kind of number one thing it offers is fantastic students and great postdocs as well. And there's absolutely no doubt that I wouldn't have had anything like the success I have had without access to these fantastic early career researchers. They're a real inspiration to work with and they surprise me every day with what they're capable of. So that's really amazing. 

[Musical interlude] 

Marianne Freiberger: Hello and welcome to Voices of Mathematics, the podcast from the mathematics faculty at the University of Cambridge. I'm Marianne Freiberger. 

Rachel Thomas: And I'm Rachel Thomas, and we're from the Outreach and Communications team here.  

[Musical interlude] 

Rachel: The person you just heard is Richard Samworth. He is Professor of Statistical Science and the Director of the Statistical Laboratory at the Department of Pure Maths and Mathematical Statistics here in Cambridge, also known as DPMMS. 

Marianne: Yeah, and Richard is indeed a very successful statistician. For example, last year he won two important prizes in statistics within 24 hours. One was the David Cox Medal for statistics and the other the Guy Medal in silver. 

Rachel: Well, that was a successful day! [Laughter] And this year, another great honour has come his way. He is an invited speaker at the International Congress of Mathematicians, better known as the ICM in our lives, which will take place in Philadelphia in the United States of America in July. It's this incredibly important congress in mathematics where they give out some of the most important prizes in maths, one being the Fields Medal, which they give to four mathematicians under the age of 40. There's also an important prize called the Abacus Medal, which is for a similarly aged mathematician who makes a contribution to computer science and mathematics. There's another number of significant prizes given out at the meeting. And being invited to give a talk at the ICM means you've done some really important work because this congress brings together thousands of mathematicians from right across mathematics right around the world. Marianne, you talked to Richard about this. And in this podcast, we'll find out what kind of work Richard is going to speak about in his ICM lecture and also how his field, statistics, is very theoretical and very applied at the same time. 

Marianne: Yeah, and I started out by asking Richard what he is most looking forward to at the ICM. 

00:03:02 Richard: I'm actually quite intrigued to see what the differences are between a big maths meeting and a big statistics conference. I've never been to a major maths conference before. So that's certainly one thing. I'm sure there'll be some cultural differences. But of course it's where the Fields Medals are handed out, so it'll be really exciting to see who's going to be awarded those. I probably won't know the people, but I'm sure there'll be a great sense of excitement around the place. 

Marianne: So what is the biggest prize in statistics then that is sort of comparable to a Fields Medal? 

Richard: So there's the COPSS President's Award [Committee of Presidents of Statistical Societies], which has the same age restrictions as the Fields Medal, but it's awarded once every year as opposed to every four years. So there's a little bit less of an accident of birth effect. 

Marianne: And could you explain what kind of, I mean, if you know already, what kind of aspect, what aspect of your work you're going to be talking about? 

00:04:04 Richard: So the title of my talk is going to be ‘Non-parametric Inference under Shape Constraints, Past, Present and Future’. So this is a topic that's very much home turf for me. It's something I've worked on for over 15 years now. But I think the field's taken a couple of really interesting directions in the last couple of years and I've been involved with some of this work. So I'm looking forward to talking about that and getting people's feedback on the work. 

00:04:30 Rachel: So this non-parametric inference, what's that all about? 

Marianne: All right, so here's a standard example from statistics. Imagine you want to understand the distribution of heights of people in the country. So that's how tall they all are. Now, it would be great if you could draw a curve which tells you for each possible value of a person's height, what proportion of people are that tall. And that curve would then also give you the probability of a person picked at random having a height that sits in a particular little interval, say between 165 [centi]meters and 170 [centi]meters. And that curve is called the density of the distribution your data came from. 

Rachel: So, now I kind of know where you're going with that. There is this curve that's called the normal distribution or the Gaussian distribution, which is that really familiar bell-shaped curve, which is given by a mathematical formula. And as long as you have collected enough data, the normal distribution will approximate the distribution of heights for the example that you're talking about really well. 

Marianne: Yeah, that's right. And it is that famous bell curve. And the sort of the highest point of the bell where you would hold it if it was a real bell, like in the middle, that corresponds to the most common height. And then as the tails slope off to the side, to the left and to the right, that corresponds to larger and smaller heights that get less and less common as they get more extreme. If you assume that the distribution of heights gives you this bell-shaped curve, which is given by that particular mathematical formula, then all you need to do to find the best fit is estimate two parameters: the mean and the standard deviation. The mean defines the highest point of the bell curve and the standard deviation defines its width. And when you do that, so you find the best fit, that's called parametric inference because what you're doing is estimating the values of parameters. Here's Richard again. 

00:06:23 Richard: So let's suppose that you're interested in estimating the density of the distribution which your data came from. One approach to this in a traditional way would be to assume beforehand that the data come from some parametric family of distributions, for instance a family of normal distributions. So that's the famous gamma distribution. 

Marianne: The bell curve. 

Richard:  The famous bell curve. And if you're willing to do that, then all you have to do is to estimate the mean and the standard deviation of the normal distribution. And you're good to go, you're done. 

00:06:54 Rachel: As Richard indicated, the normal distribution is just one family of curves, which is suitable for when you're looking at different things like people's heights, and those parameters you hone in on, the mean and the standard deviation, they set which member of that family of curves you use to describe your data. But when you're looking at different types of data, so instead of heights, for example, say you want to see how many social media posts arrive on your feed every minute, then instead you'd reach for a different family of curves. 

Marianne: Yeah, that's right. And when I learned statistics as an undergraduate, I was quite mystified by all of this because I thought like, ... how... where do all these mathematically defined families of curves come from and how do I know which one to use for a given data set? And how can I even be sure that there is such a kind of predefined mathematical formula for my data? Now that was a bit naive of me because in many cases there is a good theoretical reason for why a particular family of curves or mathematical formula would fit, but it's not always the case. And non-parametric inference addresses this. Here's Richard. 

00:08:05 Richard: The trouble is that many data sets don't follow a normal distribution well. And so the idea of non-parametric inference is that you want to be much more flexible about the class of distributions that you want to fit. So, a very simple example of a non-parametric density estimator would be a histogram. So, you just divide your real line into usually equal width bins and you count how many observations fall into each bin. And there are more sophisticated versions of histograms that you can use, but I'm particularly interested in versions of non-parametric density estimates that impose a sort of qualitative shape constraint on the density. For instance, it might be a decreasing density or something. I'm particularly interested in instances where it's log-concave, so the logarithm of the density is a concave function. 

And the reason I particularly like this area is because you kind of get the best of both worlds between the parametric and the non-parametric world. You have the convenience of working with a fully automatic estimator as if you were in the parametric world, but you have the flexibility of this infinite dimensional class of functions as if you were in the non-parametric world. 

00:09:20 Rachel: Right, so I get it. You're not restricting yourself to only reaching for one particular family of curves. But on the other hand, you're not allowing yourself to go completely wild either. 

Marianne: Yeah, that's right. And you're constraining the shape of your distribution to be one of a very large class of possibilities. And the class of log-concave functions Richard mentions actually contains the normal or the Gaussian distribution, as well as many other ones that are familiar to statisticians. 

[Musical interlude] 

00:10:04 Marianne: And so you said there have been a few developments in the last few years. Is it possible to explain them or are they too technical? 

Richard: I can have a go. 

Marianne: Yeah, let's try! 

Richard: Well, so originally much of the field worked on what you might call traditional non-parametric estimation problems like density estimation that we talked about or non-parametric regression. And I think in the last couple of years we've seen the field evolve so that shape-constrained ideas are being applied in new contexts. So, one context that I've been thinking about is doing linear regression. 

00:10:41 Marianne: Right, linear regression. Come on, Rachel, tell us what that is. 

Rachel: [Laughs] Okay, so very loosely speaking, imagine you've got two variables you think are probably related to each other. One depends on the other. So for example, ice cream sales and temperature. It makes sense to think that the higher the temperature is, the more ice cream is sold. That makes sense to me. 

So in linear regression, you assume that the relationship between these two variables can be captured by a certain type of relatively simple mathematical description. So for example, say expected ice cream sales equals some number times the temperature plus some other number. You then estimate from the data what values the parameters, the sum numbers in my mathematical description, should take to get the best fit. 

00:11:31 Marianne: Yeah, and the reason that this is called linear regression is that you assume that the relationship between the mean of the dependent variable, so that's the ice cream sales, and those parameters, that that relationship is linear. And that's a little technical, but we'll put some links into the show notes. Now you can do this for more than one variable too. For example, you could do linear regression to see how ice cream sales depend on temperature and on the amount of rain. But the data is never going to fit your neat mathematical expression exactly. There will be errors, so there will be discrepancies. And if you want to take account of those errors, you need to think about how they are distributed, how large they're going to be. So, you're back to the problem of dealing with distributions. Often, it's assumed that the errors follow a normal, so also called a Gaussian distribution, which we've met before. But let's go back to Richard. 

00:12:23 Richard: One context that I've been thinking about is doing linear regression. But where you don't want to assume that you have normal errors, normally distributed errors. So, then you can combine ideas of shape constraints together with robust statistics and information theory to provide a new and more efficient way of estimating, of doing linear regression. 

Marianne: And that's because you're applying that to the error. So you're applying that, you're no longer assuming a particular shape of the distribution of errors. 

Richard: Exactly that, yes. So, we teach our students in their early statistics courses that the Gauss-Markov theorem protects them against departures from Gaussianity for the errors because it says it's the best linear unbiased estimator. But what I've realized is that restriction to linear unbiased estimators is a very strong one and that if you allow yourself to consider other sorts of estimators you can actually do quite a lot better. 

00:13:22 Marianne: And how do you find these things out? I mean, so to what extent do you look at real data the whole time and to what extent is it purely theory? 

Richard: It depends on the project. Some of my projects are more motivated by specific data sets or application areas than others. In this case, I would say that linear regression is such a ubiquitous problem that I didn't need a particular data set to think, I know that... is linear regression an important problem or are errors always normally distributed for linear regression? I know they're not. So, this project was not so particularly motivated by one particular data set. Some of my other work maybe more closely so. 

Marianne: And what kind of data sets have been very interesting for you, like where you went like, oh, this is something new or there's something in there that hasn't been covered by theory? 

Richard:  Well, an applied data set I was working on recently was about the career intentions of UK medical students and realizing that over the last decade there's been some very significant changes, that many UK medical students are no longer intending to go into UK specialties. Many are going to Australia, New Zealand, but even leaving the medical profession afterwards. So this has pretty strong implications for education and for policy. 

Marianne: Okay, so that's interesting. So on the one hand, sometimes you do like pure theory, where it is, just theory. And sometimes you literally look at data and you analyze it and you come up with conclusions that have a direct impact on everyone. So you're doing both? 

Richard: I do both. I should say that most of my work is methodology and theory.  

[Musical interlude] 

00:15:30 Rachel: It's not surprising that Richard does so much theoretical work. It would be hard to find anyone in DPMMS's Department of Pure Maths and Mathematical Statistics who doesn't really. But I know there is a very practical component of the Stats :ab, and that is the Statistics Clinic. 

Marianne: Yeah, and that's a very interesting thing. It's open to anyone with a statistics-related element. Here's Richard. 

Richard: But then on the other hand, I've run for more than 15 years now the Statistics Clinic. So this is where once a fortnight anyone in university can come and get help with their statistical problems. And so that's exposed me to applied problems from really a huge range of application domains. 

Marianne: So apart from that one about medical students' intentions, what other kind of applications and questions do people come to you within a statistics clinic? 

Richard: A huge range and there's very few questions that you really say are standard, even though many people come along saying, I'm sure this must be very easy to you or something like that, but really you have to think from first principles a lot and I think it was David Cox who said there is no such thing as routine statistical questions, only questionable statistical routines. 

Marianne: Yeah, good quote that! [Laughter] So you get people from all departments coming. I think when I interviewed you before once you said about people coming from musicology and things like that. 

Richard: I mean, probably the majority of people are from the science side and maybe particularly the life sciences. But it has been very striking to me that the range of subject disciplines that come to the statistics clinic. 

Marianne: Yeah, that's amazing. So then do you personally go away and solve their problems or is it something that like the postdocs do and the PhD students? 

Richard: I don't go away. We do it there and then. We do. We try and help them there and then. So typical consultation lasts 45 minutes. And usually we're not going to end up on authors, as authors of the paper that they're working on. Maybe we'll get an acknowledgement or something like that. Occasionally the consultations turn into collaborations and then of course we'll be working on the problem outside the clinic times. 

Marianne: Cool. And is that something that involves younger people, early career researchers, so they get that exposure to working on people's problems? 

Richard: Exactly. It's something I really encourage my group members and other people in the Stats Lab to get involved with because I think it's really important and a good way of training you in applied statistics. I think actually, for instance, it teaches really good communication skills. One of the things that I have learned from this process that I didn't necessarily know when I started out was the way that you need to tailor methods that you would recommend to the statistical expertise of who you're speaking to. It's not realistic to imagine that everyone's going to be able to absorb all the nuances of what might be going on. So that's one of the challenges. 

00:18:49 Rachel: Aha! That's one of our favourite things to tell people when they're communicating: know your audience. So as members of the Comms team in the Faculty of Mathematics and as editors of plus.maths.org, we also teach communications training courses to researchers. And as I said, that's the most important rule in any form of communication, know your audience. You need to understand your audience for any form of communication to be effective. But, anyway, you have an interesting example of a problem that has been brought to the Stats Clinic, haven't you? 

Marianne: Yeah, this is something that Richard told me about a few years ago, where people from the Department of Music came to the Stats Clinic. And the background is that, well, basically a piece of music can be regarded as a sequence of data points. So, you know, data points is exactly what statistics deals with. And the problem that music colleges from the University of Cambridge brought to the Statistics Clinic involved finding patterns within a piece of music that would reveal the composer who had written it. So that is an advantage of being in a big university where you can make use of other departments that have expertise that you don't have. There's an interchange. But I went on to ask Richard about what it's like working in DPMMSs in particular. 

00:20:09 Marianne: So you're a member of DPMMS, the Department for Pure Mathematics and Mathematical Statistics. To what extent has being in that department, like what are the key factors that have sort of supported you and contributed to you being such a successful statistician? Like you're going to the ICM as an invited speaker and you've won lots of prizes. 

Richard: Well, thank you. DPMMS is a fantastic research environment. I think the kind of number one thing it offers is fantastic students and great postdocs as well. And there's absolutely no doubt that I wouldn't have had anything like the success I have had without access to these fantastic early career researchers. They're a real inspiration to work with and they surprise me every day with what they're capable of. So that's really amazing. The building is a lovely place to work in. I like the fact that all of Maths is on the same site and everyone's nearby like that. I think it's a little bit unfortunate that we're, for instance, so far away from the Biostatistics unit on the Medical Sciences campus, but you can't have everything right on your doorstep. It's a lovely place to work. Colleagues, the support staff, I find them a very congenial group of people to work with. 

[Musical interlude] 

00:21:52 Marianne: That brings us to the end of our interview with Richard Samworth. Rachel, you and I are going to be at the ICM in Philadelphia, and we're very much looking forward to Richard's talk. 

Rachel: We definitely are and we are definitely looking forward to hearing all of the speakers from the Maths Faculty there and of course many of the others. So, stay tuned for a podcast from the ICM! But if you enjoyed listening to this podcast, please consider recommending it to a friend or rate and review it wherever you are listening, it really helps other people find us. Thanks for listening and bye for now. 

[Musical interlude]