Description
Lily and David discuss the challenges of working with rainfall and climate data, exploring ideas of data quality, data rescue, and data accreditation. They reflect on different sources of climate data—from weather stations and satellites to reanalysis products—and examine how these can be evaluated for specific applications such as agriculture. The conversation also highlights ongoing research into rainfall intensity, satellite validation, and the importance of building evidence around which climate products are appropriate for different contexts and uses.
[00:00:07] Lily: Hello and welcome to the IDEMS podcast. I’m Lily Clements, a data scientist, and I’m here with David Stern, a founding director of IDEMS.
Hi, David.
[00:00:14] David: Hi, Lily. I’m looking forward to continuing a discussion that we started from ePICSA, but really got to elements of data quality and maybe even data accreditation.
[00:00:31] Lily: Yes, yes. So on a previous podcast we discussed ePICSA, which is, I’m gonna incorrectly say it again, Participatory Integrated Climate Services for Agriculture, and we got into a bit in there on some work we’ve done on data accreditation, on data quality, and on data rescue. I wonder if we should quickly run through what sort of things we mean by these three.
[00:00:58] David: Okay, do you wanna start that off?
[00:01:00] Lily: Great. I’ll start with data quality. So, we have this data, we have these raw data sets, and having them of a good quality is kind of dealing with erroneous values in there: it could be a typo, it could be someone meant to put in that the temperature was 20 today, but actually they wrote in 200 degrees. Or it could be that a missing value has been replaced with a zero, which as we said before, when it comes to temperature, these things are a little bit easier to pick out, 200 degrees or exactly zero degrees. But when it comes to rainfall, it requires a little bit more. As Roger says, you need to be a bit of a data detective.
[00:01:40] David: Absolutely. And should we just go into data rescue very quickly as being sort of the process which often involves going back to the paper records and digitising or re-digitising and, of course, completing these records to make sure that the digital data, the digitised data can be used for products and services.
[00:02:05] Lily: Yes, I think that that’s a good way to put it. And then data accreditation is this idea of coming up with a system of, okay, your data is of a certain standard, it’s of a good quality, you’ve gone through, you’ve looked through these erroneous values, you’ve done corrections here and there. You’re not always certain, of course, can you ever be a hundred percent? I don’t know.
[00:02:29] David: Absolutely, quality control is never finished because in some cases you just genuinely don’t know. My favourite example of this is that if you do an analysis of something like historical rainfall and you look at the day of the week with respect to the rainfall amounts, you can tell a lot about the quality of the station collectors by checking: is this evenly distributed across the seven days or is your Monday value often quite high and your Saturday and Sunday ones low, where actually rainfall over the weekend has only been measured quite often on Monday because people sometimes didn’t bother over the weekend? These sorts of things are visible in the data, and you as being a data detective, you can find this.
But to come back to the data accreditation, the data accreditation for me is really the centrepiece of this. And it’s about giving the ability to be able to say, well, we have gone through these processes, these data quality processes and we have, let’s say, stations that have at least 30 years, which is a WMO recommendation for certain analysis. And so you have at least 30 years worth of data.
That’s something where, I guess, we can set up different ways of accrediting it. I’m really excited about this because I think once we can accredit, we could also accredit for some of the other products, the satellite products, the re-analysis products. We could say, yes, this product is good enough to be used for this, in this context. Because that’s an analysis that needs to be done.
Really at the heart of Danny and John’s PhD topics has been this process to establish how we investigate the satellite and the analysis products, these remote sensing products, as proxies for the station data for use in specific applications, such as PICSA. That’s something where, again, I’m loving the idea – and I was just discussing with Danny that it’s something which we are actually well placed to start thinking about – of making it really easy not to just do research where research always has to be original, but to build evidence where you could be applying the same methods to different contexts without necessarily hugely differing original interpretations, but just building the evidence, and actually having mechanisms of that evidence building being publishable, not necessarily as research publications, but maybe as “white” or “grey” literature, or in other formats.
And the importance of that in terms of building up the evidence and using the same methods across different geographic locations or contexts is really appealing, and I think very important.
[00:05:42] Lily: Interesting. I mean, that’s definitely something I want to dig into with you. Firstly, I’m new to this world of climatic data, so I just wanted to mention – in case it is unknown to others as it was to me a few months ago – what we mean by satellite data and station data. Throughout the countries, such as throughout Zambia, you have different stations set up, they might also be volunteer, where someone is collecting the amount of rainfall and they write down how much rain they got in there each day, is my understanding.
[00:06:14] David: For example, and the synoptic stations are very interesting ’cause a lot of those are hourly data, so you actually have somebody collecting data on the hour, every hour. And then there are the three-hour data ones. There’s daily data as you mentioned. There are automatic stations that do it every five minutes. And these would all be actual met stations collecting meteorological data. And there’s a wide range of elements that are collected in different contexts.
[00:06:47] Lily: Yes. And then there’s these stations which are at certain points within the country..
[00:06:53] David: Often at airports.
[00:06:54] Lily: Often at airports, interesting.
[00:06:56] David: Because airports need this data to be able to safely land planes.
[00:07:03] Lily: Yes, interesting, I have wondered why often at airports. Okay.
[00:07:08] David: So this is all really important in terms of the aviation industry and being able to safely land planes, it’s really important. And there of course, things like wind speeds become very important elements as well as rainfall and temperature, and others.
[00:07:22] Lily: Nice. And then another way of getting your rainfall data is through satellite data, the satellites being able to see the amount of rainfall, being able to estimate, I suppose, the amount of rainfall in certain areas.
[00:07:36] David: Let me just add a little bit of detail here. Of course, the satellites, you are right to say, estimate rather than see, because if you’re in a satellite, you are above the clouds and the rain is happening below the clouds. And the clouds are really annoying in that sense that they get in the way of the satellites actually seeing the rainfall. So that’s why it’s an estimate.
And this estimate can relate to cloud cover, things which are actually measurable about the clouds. But this is imperfect. And there’s been a lot of work looking at other ways of getting data. There’s work on the infrared spectrum and actually how you can interpret that, as well to get evidence of this.
There’s a group that I’ve heard of, which I’ve not actually seen products based on their work, but I’ve heard that there was a group that investigated the NDVI index, which is a measure of greenness, and it was able to identify large storms because of the greenness that emerged afterwards.
So I think there’s some really interesting stuff, this is still a field in development of what are the sources of data. And just one other incredible source of data where there is some work happening is around the mobile phone networks. The mobile phone networks, it’s very interesting, they have to be able to deal with noise in the system and still communicate reliably. But actually, if you measure the noise in the system, well, when is the noise in the system higher? Often when it rains. You know, the water is getting in the way and that’s creating noise in the system. So there are people who have been trying to use – and I know people in Zambia who have done some work on this – the mobile phone data, the noise in the mobile phone data, to estimate rainfall. And that’s a really interesting thing.
So there’s all sorts of different possibilities and sources, not just satellite data. And the last one that I want to mention is the reanalysis data, where the best way I have of explaining this is a “forecast”. You know, if you get a weather forecast, when you want to know what the weather’s going to be like tomorrow, that uses these big climate models to try and predict what’s going to happen in the future.
A hindcast is where you use those same models to predict what happened in the past. And you can imagine that this is actually something which the models could do quite well, because they should be able to do it at least as well as they’re able to do early forecasts. And reanalysis data sets cover a whole wide range of elements and they’re available. The ERA5 hourly is one of the most well known, and ERA6 is about to appear, or maybe has just appeared, I haven’t seen the latest announcement. So that’s then another source of data.
[00:10:49] Lily: ERA6 is scheduled for late 2027.
[00:10:54] David: Okay, good. Thank you. I am excited by that because there’s potential real improvement that could happen there.
[00:11:02] Lily: Well, yes, that’s for the initial data anyway. I obviously didn’t know that off the top of my head, I did a quick lookup there. But, okay, no, this is interesting. So there’s these different sources of data. So I guess I want to go back to what we were saying about Danny and John’s PhD work, but also about some work that we’ve done, which is about comparing.
So how do we know the quality of this data? Well, we can compare it to the station data. So you can look at the satellite data, compare it to the station data of that area, and get an idea of the quality, of how close they are, of how similar those values of say, rainfall, are.
[00:11:45] David: And I think it’s really nice that you’ve brought up both John and Danny’s studies on this and the work that you’ve been doing with Emily Black from the University of Reading.
[00:11:56] Lily: Yeah.
[00:11:57] David: Because they are in many ways very similar and yet also fundamentally different. So, where should we start? Do you want to start with some similarities or some differences?
[00:12:08] Lily: Let’s start with some similarities. I personally don’t know the extent of John’s work, I know a bit about Danny’s and I know, I guess, I know what I’ve been doing as well.
[00:12:16] David: But I can fill in the blanks.
[00:12:18] Lily: The University of Reading and also with TAMSAT is comparing different products to each other, different satellite products to one another to see if this kind of new product that has come out is good, to see if it is an improvement of the current products that are out there and how it compares to the current products that are out there. And that’s been some very, very good fun work.
[00:12:43] David: Yes, and, from a scientific hypothesis perspective, there is reason why one might hope that this new merged product would actually be improving, as it’s bringing in new sources of data. Where it could be doing better, in particular, is the intensity of rainfall, which is one of the great weaknesses that has been found of the existing systems.
They’re very good on average rainfall, they’re good on temperature, they’re good on radiation, by and large, but the intensity of rain and hence things like start of the season, end of the season, these are things where identifying violent, heavy rains, – which is part of John’s work – this is difficult and the current systems don’t do well on that. Whereas the hope is that this new product might do slightly better.
I want to be cautiously optimistic about this, that we are not expecting it to suddenly do brilliantly well, but we are thinking that the additional data which is coming in, the additional sources of data, they should be enabling this product to better identify rainfall intensity in particular.
[00:14:08] Lily: Great. No, thank you. Thank you for that clarification. And then John’s work, you said that his work is more on the intensity. So what is John’s work on?
[00:14:16] David: Well, I suppose maybe what we should start with is Danny. In this particular area, he’s done multiple things, but the big paper that he wrote was about having a method to be able to do these evaluations for applications. And that’s really important because it’s all okay to do evaluations and say, these methods are good. But they’re not good for everything. And that’s fine, we don’t expect these products to be suddenly replacing the station data, although that’s what some people would like, but it’s just not realistic at the moment.
But understanding what they are good for and what they aren’t able to do yet, that’s really important. And that’s broadly where Danny’s methodology sets out how, for particular applications, you could go through this systematic method, which will enable you to do these validations for that application.
And of course, the initial application we focused on is agriculture, and particularly the agricultural products that would be of use to ePICSA, but we’re not limited to that. And within this, one of the things that John has identified when he’s applied these methods and dug in a bit further into specific cases in Ghana and in Zambia, is that actually there’s a real issue that remains, whereby these products are doing quite well on average, but the correlation is not good, as in when the heavy rain happens.
And just the identification, if we think about trying to identify categories of heavy and violent rain, it is not something which these systems are able to do well at the moment. The one exception to that, which was quite interesting, was ENACTS and there were some interesting questions because ENACTS actually has a method of using and integrating station data into what is then displayed as this gridded product. So, where ENACTS existed and was available, as it was in Zambia, we did a bit of a study on that and were very interested to see how ENACTS came out. And that’s actually a further study that we hope he’s going to do much more, but it’s a challenging piece of work.
[00:16:55] Lily: Yes, I bet. But it sounds fascinating.
Okay, so should we then go into the similarities between these three areas?
[00:17:01] David: I suppose all of them are looking at comparing the satellite products, the analysis products, to each other and to the actual station data. All of them were doing so in some sense from a point to pixel or pixel to point perspective, which means that you’re not really comparing like with like, but this is exactly where the methods paper is detailing how this can be done and why should this be done and what’s the importance of doing this. A lot of Danny’s paper was on that method.
But the big difference is that in Danny and John’s work, there is this strong desire to make these evaluations useful and relevant for specific applications. Whereas in the work you are doing with Emily Black, this is all about evaluating a new product. Both of these are useful and both of these are needed, but they are different in nature in terms of what you’re trying to do.
That’s not a bad place to finish actually on these similarities and differences. And maybe it’s worth just tying this back to the idea of accreditation, that it’s all well and good to sort of have an accreditation of a product saying “this product is reasonable, it’s good”.
But what I think Danny and John’s work is highlighting is the more useful accreditation would be a way of saying that this data is appropriate and good for this specific application, and that that specificity will give different results and will add value to the studies to have that specificity associated.
[00:18:59] Lily: Nice. No. Excellent. Well, thank you very much for outlining that.
[00:19:04] David: No, this has been great. Thank you.
[00:19:06] Lily: Thank you.

