07. Why Better Data Matters

The IDEMS Podcast: Alternatives to AI Empires
The IDEMS Podcast: Alternatives to AI Empires
07. Why Better Data Matters
Loading
/

Description

Continuing their examination of the assumptions underlying today’s dominant AI narrative, David and Kate explore what makes data useful, trustworthy, and meaningful. They discuss the limitations of extraction-based approaches to AI, the importance of local context and data ownership, and the challenges of building systems that can learn across diverse communities without centralising control. The conversation highlights why better data—not just more data—may be key to building more effective and trustworthy AI systems.

[00:00:07] David: Hi, and welcome to the IDEMS podcast. I’m David Stern, a founding director of IDEMS, and it’s my pleasure to be here with Kate Fleming again, another fellow director.

Hi, Kate.

[00:00:16] Kate: Morning, David.

[00:00:18] David: Morning. So what’s the plan for today?

[00:00:22] Kate: Continuing our series, which emerged from your Earthkeepers versus AI Empires event you attended, and Karen Hao’s book, the Empire of AI. So we are continuing to be loosely in conversation with the book, I would not say we are trying to really stick to her line. But the topic that I think is really interesting for us that we have glancingly touched on, but haven’t really explored is high quality data.

So, I would say, if I had to divide that into subsections of what I’d like to talk about, it’s current approaches to collecting data and is that high quality, barriers to collecting high quality data, and then our sociotechnical vision for that. And maybe within that, it’s also even worth saying, well, what does high quality data even look like, how are we defining that?

And so I’ll set up that as the frame for this conversation, and then there’s just one quote which I will pull from Karen’s book, which I think is an interesting window into current approaches and the models that it sets up. So there’s a quote from Ilya Sutskever, who’s the co-founder of OpenAI. He’s no longer with OpenAI, he started his own new thing. But this is from late 2024, so this is relatively recent, and he gives this quote: “While compute is growing through better hardware, better algorithms, and larger clusters, the data is not growing because we have but one internet”. And so, I’ve just thrown a lot on the table.

[00:01:54] David: No, no, no, the place you’re starting, “we have but one internet”, is a perfect place to start because Lily and I did an episode quite a while ago when she did her ankle in, where she then went to hospital and she was trying to figure out, she was bouldering and she was trying to figure out how high the wall was from which she fell. And she got the AI response, and the AI’s response said that the centre she was at had just built a 25 foot bouldering wall.

[00:02:29] Kate: Didn’t she say 25 metres? I thought it was something that was insane.

[00:02:32] David: It was insane, whether it’s 25 feet or 25 metres, either way it’s not a bouldering wall. And she thought that’s rather high, this doesn’t sound right, and had great fun tracking it down to the first post from the centre as an April fool joke, which of course the scraping of the internet couldn’t distinguish from the real information and was now out there forevermore as the height of the bouldering wall, which AI would pick up.

So this is, you know, we have but one internet and that internet is not necessarily accurate, although it was correctly picking up information that was posted on the web at some point about that centre, but it was posted as a joke, it was an April fools joke. And yet, of course, it couldn’t distinguish between that and a real post announcing the real height of the wall.

And I love that example. And that is exactly what you would expect from a system which believes the internet is the best source of information. It is a wide expansive source of the information, once you’ve started using the whole of the internet, it is pretty clear that anything else looks small in comparison. And so if your business model is based on more compute and more data, yeah, there’s only one internet and you can’t get much more information that you’ve got. But the internet is not what I would consider high quality data.

And this comes back to the sort of core point that the real advances are not going to come from more compute, more data, they’re going to come from better data and more accurate data, high quality data as you put it, where there are elements of trustworthiness built in.

I loved the fact that basically the best source of the internet was broadly Wikipedia for a long time because there are quality controls in that. That’s not to say Wikipedia is perfect, but if you are taking the internet as your source of information, well, Wikipedia is not a bad place to start on that, as the foundational source of data because it has procedures set up, it has processes in place, which when there are disagreements they get resolved in certain ways. That doesn’t mean it’s perfect, but it is more reliable than many other sources.

[00:05:22] Kate: So there are so many questions I could ask out of everything you just said, but I think underlying this and the fact that you’re pinpointing Wikipedia is you have to look at the existing software platforms, the existing things that are collecting data, how they’re designed. And so this does get into, well, what are current approaches? So how are things designed?

And there’s the public internet, publicly available, but then you also have obviously data that’s within companies. But those systems are often designed to be surveillance, I know can sound like a very loaded term, but they are not multi-directional, they are looking into what’s happening and trying to just pull that data out without having to do any, well, I’m not even sure how to frame this question exactly, but yeah, I don’t know, do you hear what I’m trying to… 

[00:06:26] David: What I’m hearing is two different things. I’m hearing that if we think about how AI is currently built, it is sucking data out from anywhere it finds it, and it’s a one dimensional process where you don’t have those easy returns of information, which check the validity of the data.

If they had gone back, if they had a process to go back and whereby that centre could say, no, that was an April fool’s joke, please ignore this when you are thinking about what the height of our walls are, then that feedback loop would’ve been able to resolve the problem that Lily had with the data that was pulled from the internet. But it’s a one-way process where those quality measures, those quality control pieces, there’s no real mechanism for that. So that’s one aspect that I heard.

[00:07:16] Kate: Well, and just to jump in there, I wanna save for another episode the role of humans in building AI, but there is a lot of, and I will use quotes around quality control, I mean, that’s what reinforcement learning through human feedback is supposed to be doing, you are trying to have humans clean up this data and it’s just that some data might not get flagged as data that even needs to be checked or needs some sort of human feedback, it’s just accepted as fact. But yeah. Okay, so keep going.

[00:07:44] David: Well, this is the thing where if you can, and you can have feedback things and there have been AI processes which asked for feedback, and so you could give reinforcement learning. If enough people got that post and gave the feedback, this is not right, then that reinforcement learning could eventually learn that that isn’t the right height of the wall.

 These processes can exist within it, but they’re not systematic in terms of the ownership because the ownership of where that information’s come from, and therefore the influence it’s having, is not something where there’s a feedback loop in that way. There’s a disassociation between the people giving feedback and the people who actually own the data which is being exploited. And that’s a very deliberate choice, this is this extraction process.

Whereas what you are alluding to, I felt in what you were saying before, is the fact that, really, if it wasn’t a one-way extraction which is then processed, that would be quite different. And that comes back to your example of a company and the data that they hold and the efforts that are happening towards company brains and this sort of thing, where you actually are trying to use that data to develop the company processes and to have AI systems working on company level data.

For big companies there could be sufficient data to actually develop things like company culture with the right feedback, reinforcement learning, where these things could be built out in ways which are being investigated right now. But what I want to come back to is the key point. What is lacking at the moment, I would argue, is that even in the company brain processes, they are looking at an extraction process first, and then a reinforcement process. The concept of data ownership, and actually who owns the data and the quality of that data as it’s getting interpreted, that’s the bit that’s missing.

And my belief is we cannot get really high quality data through a purely extracted process because you’ve lost the context relevant knowledge. And to me, coming back to the data quality question, the key failing that we have with our current implementations of AI is that they are built on extraction. And once you’ve extracted, you’ve lost the ability to really exploit the contextual information that would’ve been available before extraction, if you stayed within data ownership, if you want, as part of the models you’re building. 

[00:10:28] Kate: And I think part of what you’re getting, I mean, so much of this is structures and logic of power and how systems should work. You referenced that company brain, which is, I think, an emerging field where it’s like AI as this thing that can make sense of these messy systems that are quite disjointed, where different humans along with different technologies hold different things.

And I think there’s this conceptualization that this is an outdated model that goes back to organising armies, you know, whatever the metaphor is. But you just have to use this telephone system of information bubbling up because it’s the best we can do with what we have. And so it’s only perceived as a weakness, not the idea that in breaking things down, distributing power, it’s because what you recognise is that that team on the ground in some place actually needs some autonomy because they have variables at play that you simply cannot understand or cannot see at the top.

So if you are trying to build this system where your assumption is all you’re trying to do is suck out that low level, whatever they hold, that’s not actually that valuable, that’s mostly an impediment, but they’re kind of doing something, we could just turn this into a system and AI can just do this.

I think the point is you’re missing a whole lot of data, and data I use broadly there, it’s knowledge, it’s contextual understanding, whatever it is, there’s a lot more at play there than this kind of reductive view of you’re just trying to pull data out and systematise it.

[00:12:07] David: Absolutely agreed. And I think the thing which this brings out for me is exactly this idea that, well, what would it look like if AI was embedded on more fragmented data ownership, actually where the data that’s owned and where the systems that are being built, the reinforcement learning is coming within context rather than across context.

I feel that there’s a real opportunity here for there to be systems which are actually serving local communities much better, but also serving a global community because they are contextualisable. And so I think there’s a real possibility for us to use the methods of artificial intelligence as they are currently, as they have been developed, but in ways which are really rather different and new.

And there have been some instances of this, but they’re few and far between, and they require a different logic to the “more data, more compute”.

[00:13:17] Kate: So in some sense what I’m hearing is, if we take that company brain metaphor, it’s the difference between that kind of thinking, which is you’re just trying to get it all up into the all knowing, all seeing Oz and that’s where all the sense will be made by this AI brain, and it’s much more about a lot of smaller AI human systems.

And so we’re not even to the point where we might begin to think of how those might converge. We’re getting so far ahead of ourselves. We need to just be building these much, much smaller, more constrained systems. This gets back to that expert idea, but within these very narrow use cases. That’s where you actually start to get at very high quality, very meaningful data.

And then there will be another layer of very hard work, which is how do you bring those into interoperability? How might you start to navigate sense making across all of that distributed system? But that is a future state problem, which we’re not even ready to address. Is that kind of a fair encapsulation?

[00:14:21] David: Yes, but I don’t think we can shy away from the fact that actually the idea of the company brain is one where, done right, it is what we’re describing. A company brain which doesn’t do this, if you have a big global, multinational company and they don’t take local context into account in the right ways, this is going to be a disaster.

And so, the only point I disagree with you about is the fact that we have to sort of do the small bit before we actually tackle the harder problem of bringing it together. I have a feeling we need to just get on and try and do this because it’s needed everywhere. Yes, the small pieces are correct, this is what we need to be doing. We need to be moving towards small language models, we need to be moving towards these sort of specialist AI agents rather than single big artificial general intelligence.

But pulling it together to create a company brain or an organisational brain, which is actually effective, that’s the same problem that we need in the social impact space as it is in a company. So I guess my only disagreement is on the level of ambition, I’m that bit more ambitious. I think we have to take on that bigger problem.

[00:15:42] Kate: Yeah, and I agree that that is what we constantly see an impact on. And I will say we’ve been using the company space as the metaphor just because that’s where existing AI is focused, because it’s where the market is, it’s where the money is. But when you look at the impact space, we have never held that thinking independently because there is so much interplay between what happens in a local community and what happens across a system.

If you are not holding both of those at the same time, you are either missing the trees for the forest or missing the forest for the trees, you’re not getting the whole picture, and so you’re not going to drive the same impact. And so it gets back to, it’s not necessarily hugely efficient early on, but it’s very effective as you’re trying to hold space for all of these elements.

[00:16:36] David: No, absolutely, you’re spot on here. The company brain analogy is useful for us because a lot of people understand the company’s perspectives. But most companies want to centralise. This is a key driving fact, the company brain, as it’s currently implemented and conceived, that’s what it’s doing.

Whereas nobody is trying to centralise across diverse contexts in the impact space because everybody knows that this is something where that diversity of marginalised communities, of different groups, of context is so important. So the idea of just centralising all of that is something which there are a few people in international development or social impact that take that view, but they’re in the minority, whereas the majority recognise the importance of context and working locally and working in different contexts.

And so that’s where what we’re suggesting and putting forward is so natural and important in the impact space. I believe it applies to the company space as well, but I can see how the attraction in a company saying, no, we can centralise all this and hope it’ll work. And maybe in some cases it will.

[00:17:51] Kate: I also think in companies – it’s part of why I don’t like working for big companies, or was never even very good at it – is they don’t actually want their employees often to have agency, to have independence. When they spin off, they are spinning off, like it is not towing the company line. And so that ability to constrain rogue employees might be a feature, not a bug.

[00:18:18] David: Yeah.

[00:18:19] Kate: And even the idea that they’re going rogue when they’re doing something that they see needs to happen, that’s just a different mindset.

[00:18:25] David: Absolutely. We’re very comfortable with the fact that we are in the right space, working for social impact, prioritising that, what we are proposing as this alternative way of doing this it’s not in real debate, there’s very few people who would argue against the importance of real localization, taking local context into account, allowing local knowledge to be able to drive local processes. These are things which are deeply embedded in the cultures that work in social impact. So we are in the right space trying to do that.

I do believe though, that this is needed across the board.

[00:19:07] Kate: Yes, I think we agree that it is something that can have value. I just think making the business case in the immediate is a harder thing to do for us, and also not where we’re interested in working. Fundamentally, it’s not the problems we’re interested in addressing. So, you know, that’s a big piece of it.

I guess one thing I would also say here is, we’re very focused in this conversation on AI, but something we are very aware of – and bless our mathematical team for holding the brains to be able to conceptualise this – is that this opens up another layer of challenge, which is data ownership, data interoperability, how you build the underlying system with existing or new open source software, different solutions. How are you enabling data that otherwise would sit in isolated, kind of forked, whatever you wanna call it, kind of spun out into total independence, how are you keeping that in conversation so you are getting the big data, whatever the term is, but that very massive pool of data needed, but you’re doing it in this totally different way.

[00:20:17] David: There’s of course subtleties in there, but I do want to just call out the work that has happened, particularly in the medical field around this, where data sensitivities have meant that sharing the data is not an option because it’s highly personal data. However, the methods that are being built are showing that you don’t need to share the data to have this deeper learning.

And actually the deeper learning which can happen without sharing the data, builds models which are not violating anyone’s privacy, but actually getting learning at scale on higher quality data because the data is not being shared and therefore it’s safe for people to have high quality data in other ways. So I think there are models of this, which we are very well aware of, and which we are trying to learn from and expand on.

[00:21:14] Kate: I’m just gonna interrupt you there. Could you talk through – because you just said some things that seem to be in conflict – they’re not sharing the data, but they’re learning on high quality data. Will you talk through what that scarcity sounds like, or the data protection approach. 

[00:21:29] David: The key point is, you know, we have to remember that AI is not magic, it’s just mathematics. So I’m a mathematician, everything’s mathematics. But the point is, what reinforcement learning is doing is this is improving models based on the data which is fed in. Now, the point which has been articulated really well in these medical fields is, well, you don’t need to share the data to get the benefits from the improved models. Instead of sharing the data, you can share the models.

So you can send the models into the server, which holds the data, which can then do the learning, the reinforcement learning, on that data, and then instead of sharing the data back out, you share the improved models back out. And so you are getting the improvements that come from using that data into the models without actually sharing the data, by sharing the underlying mathematics or the underlying models.

And this is an obvious thing to do. We don’t need all the data in the world to be open and shareable to be able to have these really powerful learning mechanisms that require lots of data. And this is where you can have data ownership really preserved while also it feeding into these really powerful models that take advantage of data from multiple sources.

[00:22:59] Kate: So a follow up question. That was a great answer and that makes a lot of sense. So what I hear is the models are open, the models are able to work with the data, but how is that data? Because shouldn’t the data also be improving the model, you’re gathering some piece of information that makes you realise, oh, this model didn’t hold this variable that it needed to hold. So how does that get negotiated? 

[00:23:22] David: This is where, of course there’s layers to this. And a lot of people don’t realise this, but one of the easiest ways to think about it is the fact that you have a model structure and then you have a model parameterization. So often what is actually happening is that you build your model structure and then the reinforcement learning relates more to the parameterization. Now, this is a real oversimplification because in the highly complex modelling spaces, that’s not necessarily what’s happening. But it is happening more often than you’d think.

I’ll just give a really simple example. When I work with data scientists who are PhD students who are doing this, they use the modelling systems that they are used to using, without recognising that they need to also think about these underlying model structures because the parameterization isn’t enough.

So this is something which is very important in how we use AI, how we build AI systems, to distinguish between the parameterization of the models, which is generally what the reinforcement learning is changing, and the underlying structure of the models that are being parameterized. And the fact that both of these can change is important.

And your idea of actually adding additional variables, this would be changing the underlying structure of the model rather than just the parameterizations of the models as a way of framing that. And the important thing is there is no restriction to what you can do within a context. You can change the parameterizations, but you can also do analyses, which then come back and propose a totally different model. And that could then feed back into these systems.

So the analyses that you can do on the local data that you have, they’re not limited, there’s no conceptual limit to what you can do within context versus across context. You can change the underlying structures of the model as much as you want, and then the reconciliation process is harder, of course, it’s much easier if all you’ve done is change parameters. It’s harder if what you are doing is actually also changing the structure of the models. But it’s possible.

So this is something then, I would argue, where a lot of the work could be happening. It is to understand, well, what are the different things we actually want to do within context versus across context if we are not sharing the data? But it becomes, you know, I think it really does become powerful and very valuable to do so.

[00:26:03] Kate: Yes, I think that focus on constrained context and the fact that there’s so much to be unlocked there is very useful. I do quickly spin up into, well, what if these cancer researchers hold this cancer information, but this community group holds a lot of demographic information, or the way the community works, or something like that. And so I actually need to put these two things in complement to produce this, whatever the thing is, the application of this cancer research to this particular context.

What I’m hearing partly is also – you can correct me here if I’m wrong – but this sounds like it’s more like a governance and organisation social problem than it is a technology problem.

[00:26:45] David: There is both, of course. And you’ve been talking to people like George because exactly the work that he’s been working on is to look at how we can actually take these different models and modelling systems and actually get them to work together, and that’s a hard problem. This is a technically hard problem, there’s a really interesting group in Germany who I still haven’t met who are working on this in another way. And there’s us working on this in a particular way. There’s a few other groups working on this. It’s a hard problem.

That idea of actually being able to take your cancer researchers, as you put it, who are modelling with the data they have in this particular way, and then community groups who are modelling in a different way, and to be able to understand how to bring those models together into a coherence, these are hard problems, and this is really where AI should be helping. But not the AI of just generative AI. This is, again, this issue which has come up related to the empires of AI, where the language has been captured.

I think maybe this is actually a good place to finish this because the point which is so important is that if we recognise that it is not just more data and more compute, because that’s the only vision of AI that we have, and we recognise that better data with less compute can do certain specialist things, then what we should be putting our efforts into is a different sort of AI, which just looks very different. And that’s sort of where I feel this has taken us to.

[00:28:25] Kate: So interesting. Thank you, David.

[00:28:27] David: No, thank you.