Description
Continuing their discussion on the future of AI, David and Kate explore how advances in large language models could enable a new generation of smaller, more specialised AI systems. They discuss why the next wave of innovation may come from building tools that are more efficient, focused, and responsive to real-world needs rather than simply pursuing ever-larger models.
[00:00:07] David: Hi, and welcome to the IDEMS Podcast. I’m David Stern, a founding director of IDEMS, and it’s my pleasure to be here today with Kate Fleming, another director.
Kate, we are keeping going with our AI series.
[00:00:20] Kate: We are, I have so many questions, I think our team has so many questions and we work in this space to some extent. And so I think if we have questions, it’s not hard to imagine other people have questions. And I will also say, you have been talking about these things for a long time and I sometimes just looked at you and kind of like, okay..
Last episode we talked about Symbolist versus Connectionist AI, large language models, and at the end of the last episode, I think, you actually mentioned large language models and small language models, expert AI, generalist AI, all these different things, and I think you also want to point out that these are not totally clear boundaries.
I’m not even sure where to start, but I guess what I would want to really explore in this is the relationship between those that have been positioned as camps, these two different views of what AI can be, and the fact that they’re not binary. They’re actually quite intertwined, is what I’m hearing.
So I’m gonna give the floor to you mostly.
[00:01:27] David: Well, they were camps 70 years ago. So this is the thing, they were different views or visions of how AI could be developed. And, arguably, both have developed, but one of them has gained popularity and resources way beyond. The idea that AI is about learning and the importance of neural networks in this is something which has been extremely powerful for over 30 years.
The big turning moment there was when AI won at chess, and that was neural networks. So this is, if you want, this idea that this wasn’t a rule-based system that somebody built, it wasn’t about an expert learning how to play chess better than the world champions. That would’ve been actually a fruitless endeavour. For that, learning and the learning approach and the neural network approach to learning was exactly what was required.
And this led to incredible advances in all sorts of domains. And that was what has been pushed to its extreme leading to the breakthroughs with the large language models where you put huge amounts of data in, and using that and using large amounts of computes, you then are able to get generative AI and to generate things which are incredibly impressive.
And this is, again, this is the advances and what is so powerful about the connectionist approach where we’re actually building learning, where we’re building learning cycles, these reinforcement learning cycles, which are so powerful. So that is what people now know of as AI, and people equate AI to the reinforcement learning approaches or to the products of the reinforcement learning approaches.
[00:03:37] Kate: And as the foundations for where we think, and especially relating to your conference in Zambia, the Earthkeepers versus AI empires, that introduces a lot of problems. And it’s also not necessarily getting better at solving certain problems. So I guess this is where it’s helpful to think about our perspective.
We work in impact spaces, it’s not only delivering social impact, but also thinking about what the consequences are for the communities. It’s not that you’re causing harm as you’re creating benefits, you’re trying to have a general do-no-harm sensibility. And so maybe that is the starting point for just the fact that we have a particular lens that maybe shapes our sense of what AI should be or should do.
[00:04:29] David: It’s interesting. That is a lens we take, and that does affect how we see a whole range of different things. But the lens we have on AI is actually really simple. This is just generally about where the next steps forwards are coming from. So, if we think about most things in the world, when things are balanced, that’s when they become most powerful.
At the moment, this connectionist approach has grown way beyond where we’ve got to in the symbolist approach. It’s not that there hasn’t been real progress there. The mathematical modelling approaches and things which have been understood and built there in the symbolist approach are also incredibly impressive.
But they’re not unified. They’re bitty. They’re here and there and many of them are out of date. The big crop models, they haven’t really brought in the mathematics, new mathematics, since the sixties or seventies. And so, the structures behind those haven’t been advanced in the same way. There hasn’t been the same effort and the same academic human effort which has gone into them.
And certainly there hasn’t been that effort unifying them, for good reasons. We mentioned in previous episodes that this relates to the finances, this relates to the potential to commercialise it, the need for collaboration around it, and so there is a whole range of reasons why this hasn’t happened.
But if we recognise that the balance is where the really powerful things can emerge, I would argue that the most obvious next big advances are going to be coming from actually balancing the advances, the recent advances made on the connectionist side with further advances on the symbolist side and bringing those together to build these combined systems.
That’s really a low hanging fruit, if you want, compared to pushing further ahead purely on the connectionist approach. And furthermore, actually, the recent advances that have happened with large language models, with these incredible ability to generate images and videos and audios, well, really the natural next step there, is to be able to do more with less, not to do more with more. To be able to get equivalent results with less data, with less input, which means that we can get things which are much better in specialist contexts.
So we can actually use the advances that we’ve made to build specialist tools in particular contexts, that’s where the obvious next steps of improvement should come from a scientific perspective, that is what’s most likely. It is possible that there will be advances which are coming from elsewhere, and sort of really step changes towards artificial general intelligence, which are further improvements. But it’s more likely that we actually now consolidate what we have gained, the gains we’ve made, and we are able to do them more effectively, more efficiently.
In the cycle of how things have been changing, that’s the more likely next step. It’s maybe not as exciting, but it is certainly more useful. And truth be told, when we are looking at this from perspectives of everything from the social impact spaces, but also to business perspectives, this is what people want, people want the useful applications of this to be really made real, and to be leveraged into their hands, whatever that may mean.
[00:08:22] Kate: What does that look like in practice to develop that? Who holds that work? How does it get done? Is it just that OpenAI builds its application lab that does this work? Or is it something that needs to look very, very different?
[00:08:40] David: So it’s interesting that, by and large, it’s the startup scene which is really coming up with these applications, trying to find the uses and make the uses work. But they’re getting caught up in the buzz around moving forward towards artificial general intelligence, towards the latest models.
I don’t know at this point whether the startup scene is even the right place to do this right now, because they’re not looking beyond the big tech ecosystem. I think that there needs to be, and there is, there are people who are doing research work, there are people who are doing wonderful small scale applications, but we really need a large scale effort of competent people just saying, okay, what are the useful problems that we’re really trying to solve with the technologies as they’ve now emerged?
So much more effort academically and also in industry on small language models and on how they could actually do the job of the large language models with less energy, much more cost effectively, much more efficiently, that’s something which should be a good business plan because being more efficient and more cost effective is good for business.
But the other piece, I think, which is critical is to actually have academic research freed from the driving force that is the current big tech priorities, to have more of that public money going towards the research, which will not necessarily lead to the next artificial general intelligence, but that will actually lead to the advances that make the current systems really useful for society.
So I think that there’s potential there. I think there’s also real potential in the space that we hold, which is, well, in low resource environments we can’t get lost in a tutor for kids which is extremely resource intensive. But we can look at how education systems could be built, which have similar functionality, but are much more accessible and are available on low resource devices.
So I think that there’s a whole range of places which can be advancing this alternative approach. And who knows, maybe one day it will actually become the dominant approach if the current generalist AI cycles bust. You know, there is a boom and bust cycle, which seems to be happening. If that bubble bursts, then what will be left is the people who are doing the work that really has value and adds value in context.
So, who knows? It might be that it doesn’t burst because they actually do find a further advance. Great. But that doesn’t stop the fact that the work that will have happened to actually make the current advances really useful will still have great value and great societal value, and maybe it will then be able to leverage whatever the next transformational findings are at that level.
My personal view is that I think the probability of them finding that magic nugget before it bursts, I think is low. But what do I know? And so this is something where there is a possibility.
[00:12:32] Kate: I find myself with a few questions listening to you talk. And I guess one thing that just comes to mind that’s interesting in all this is how much policy capture has been done by AI. It affects where the funding that should be going for R&D in really new directions has been kind of co-opted by, “well, how are we just advancing this existing model?”
That’s what it feels like to me when I look at funding calls related to AI. They don’t always feel like they are pushing in as new directions as they might. But I guess I have a question related to open science. I mean, one of the criticisms of OpenAI, from existing technology companies – well actually Meta surprisingly, I think with Llama, which is totally open source – is that OpenAI has been quite closed in their models, there’s sort of this antithetical to open science.
And it seems like there is a push that would serve existing, entrenched AI, to now get everyone else to produce open science data sets. And so that seems good, obviously, you want open science data sets. But I guess I’m sort of curious, those data sets alone with existing AI is not enough. That’s not enough is what I hear in what you’re saying.
My question is really: what is the right combination? Is it data plus new models? What does that look like? That’s a hard thing to answer.
[00:13:58] David: As you say, the question of whether open science, and open source even, is or has been captured is an open question that I don’t have answers to. I was a huge, huge advocate of the open movement in all its forms, and I have to recognise that there is a real possibility that that effort has simply been captured at the expense of the people who were believing and actually had good faith in it.
The legal systems are not holding this up right. It is something where there are people who are correctly really angry at how this is playing out. So I don’t have answers there. What I do know is that from a simple collaborative perspective, if we are looking at trying to build things that work well together, then open is still one of the best tools for collaboration that I know of, both in terms of open science and in terms of open source and open educational resources. This is a tool which has enabled and can enable collaboration.
Whether it also enables a form of corporate capture, that is also possible. Whether that can be fought against and how best to fight against that, I don’t know. And yeah, there are a lot of open questions there. So I don’t have answers to that. This is really part of where the world we are living in is a complex one.
[00:15:46] Kate: Yeah. I mean, even I just finished another book, a different book, an audio book, and at the end the author says, no portion of this book can be used to train AI models. And I thought, well, how do you enforce that? It’s a bit of that issue too, where it’s like you can put all these licences on, but you have such disproportionate power dynamics and access to legal systems and lawyers and teams and all these things that it can seem quite naive.
I can see why people are thinking, why would I put something in the open domain even with licensing attached to it? Because I can’t enforce that licence, I can’t enforce that copyright, it’s just going to be me versus this team of lawyers. So I can see the things that are happening that make that hard.
So I guess we don’t have the answers to that. But I think something you do have the answer to that I’m very curious about, is if the existing model of AI is just this voracious data consuming machine that then needs so much computing power, so much energy – it’s why we have these data centres, it’s why there are all these issues – or if the vision is for something hybrid.
What is the hybrid thing that is not just going bigger where you’re layering something on the existing, but it just means more? Because what I hear is implicit in what you’re saying is there’s actually a step down that’s happening.
And so I’m curious for you to talk about what does that transition step down look like? What is the role of LLMs, probably early on to get us started, and then how does that transition to a more resource efficient knowledge expert SLM AI?
[00:17:35] David: So the step down from large language models to small language models is really about, well, how can we do more with less? And this is where I think really interesting research is happening on saying, well, actually, if you take a specialist domain, you can get responses which are equivalently good or similarly good with a much, much simpler model, with a small language model rather than a large language model?
And that small language model can be fed with a lot less data, but more precise data, maybe about a particular topic you’re wanting it to be trained on. And it can still have the same interactive capabilities, but on a narrower range of topics, if you want, than your large language models.
The idea that you go to the same language model for everything, well, that’s inherently inefficient. If I’m thinking about this very simply, if I have my maths homework and I want a student to interact with a language model to help them with their maths homework, of course, I’d much rather them interact with a specialist language model than a large language model which is good at everything, because a model which is good at everything doesn’t have baked into it the things that I would care about for my small language model, which would help with that specific task.
The same would be true for a language model related to somebody interacting with a parent on their parenting. I want people to be able to just go to their language model and ask the question as they’d want to do, and as they’re doing now with the big large language models. But we should be able to get those interactions happening with equivalent, or almost equivalent accuracy and ability, but with much, much less data and compute behind by reducing the scope.
[00:19:45] Kate: What is the role of existing AI? And sorry if you sort of touched on this and I didn’t quite register it, but really, what is the role of existing AI? This wouldn’t be possible without AI getting to where it is.
[00:20:01] David: Absolutely.
[00:20:02] Kate: And I think part of my question is: there is some dependency there. How are you breaking free of dependency? where do we have to not be naive and recognise like, yeah, there is some period of time where we are going to be working very closely with existing AI?
[00:20:19] David: If you currently build an agent within the current large language model systems, then you are building your training data set, and then you are using it with the large language model systems to get a customised agent for a specific task. Well, that process is exactly the same as the process you would do if the underlying model you were using was a small language model.
Now, the big difference is the amount of data that’s trained on it, the amount of compute which would be needed for it. But also, the big difference is that you can do the large language model training right now, whereas to actually set up that system and for it to work really well with small language models, there’s research to be done. There’s actual work to be done to get this really good.
So there’s further work to be able to make sure that we can do more with less. That’s exactly where there was the opportunity for work to be done. But for us to be able to know what we’re aiming at, well, we’re really lucky that with the large language models as they currently exist, we can build these agentic AI systems which have these specialist roles and then we can benchmark the small language models against the large language models. And only once we’ve actually got the small language models that are working well enough for individual agentic AI use cases, do we then switch over?
This is the sort of thing that it might be at first, there’s almost none where we’re ready to switch. And, of course, when we start working on those small language models, what we would probably find is it’s not that one piece of work on the small language models will suddenly replace all of them. No, it will replace them in certain contexts and then we’ll have to make different advances for different contexts because the needs will be different, because maybe in certain contexts the language elements are more important or less important, maybe alternative languages are more important or less important.
So there’s so many unknowns and there’s so many ways that when you sort of cut down to small language models. There’s so many places you can build efficiencies in that you could choose different priorities and they might work in some contexts, but not in others. So we’re likely to have an explosion of different possibilities, which then play off against each other in different contexts, and eventually you get some which work and serve the different purposes you need.
So that’s what I would imagine happening over the next few years if people go down this route, it’s a choice.
[00:23:03] Kate: And this is an IDEMS strategy question, which you don’t know a correct answer to, but as you think about this, and as I think about what we work in, I see, well, the only way that this would make sense for us to work in this is really in our very specific areas of work. We are working on parenting, we’re working on biodiversity, whatever the issue is that we are working on, we are very focused.
But would you ever be interested in becoming more expert and really kind of the go-to expert in what that looks like to build, probably for impact, because that is always our area of interest, what does it look like to do this well for low resource communities, what is the sort of sociotechnical process for doing that work? Is that something that is interesting?
[00:23:52] David: I’m a mathematician, I’m interested in it even for non-impact spaces. There is a really interesting mathematical challenge. You know, I don’t often get to get into the deep math challenges anymore. Yes, I’d be fascinated to sort of lose myself in a problem like that and actually take a team with me down that rabbit hole and say, can we actually do this?
Do I think we are best placed to do that? I don’t know. Because there are people who are doing this, NVIDIA for example. There was one of the researchers in NVIDIA who wrote an article released on the arXiv recently about small language models, where they’re thinking about this. So we are not alone thinking about this. Are we best placed to do this? Well, the NVIDIA guys are much better funded than us, they’re thinking about this and they’re doing that piece well. They get to have the fun, not me, which is a shame.
[00:24:43] Kate: Well, except that I think that yes, they can keep doing it, there’s room for everybody, but I do think we are often sitting in a very special position. We have access to underserved, marginalised, different communities that I think are incredibly distrustful of AI.
The process of bringing people into participation – and I’ll save this for another episode – but I think one of the things that Karen highlights so well is that AI is not a product, it is a system. We talk all the time about the system because we are externalising the harms, the value, all of those things in the system. And when we talk about that system, we can often seem insane, like, why are you thinking about all of those things, where’s your product?
But I think existing AI has been very good at hiding the system. I know I keep bringing this up, but I guess it just makes me think about, does having access in different ways to different parts of the system and setting up a different relationship position us to build things and do things differently, or better, or more quickly? I don’t know.
[00:25:57] David: I honestly believe yes, if we can figure out how to get the growth. Because the point is really going into this space, the speed at which AI companies are growing in this is just insane. I heard there was this Cambridge startup, which literally I heard about just last week, which just got given something like 87 million – I’ve forgotten exactly what it was as a startup – just to get started.
We can’t compete with that because we are trying to build something which is slow, patient, and actually growing, but in a way which is sustainable. If somebody threw that amount of money at us, you know, I’d be scared. And yet that’s what is needed for us to actually compete.
And that wouldn’t be enough necessarily to see things over the line. That would see us to get to the starting line of really competing. That speed of growth is something which is so destabilising, we might lose our links to the grassroots, and so on.
But that’s probably the speed at which we would need to grow to be able to really enter into that lane and be the ones making it happen. And so, much as I’d love the intellectual challenge, from a business perspective, I don’t know it would be healthy for us.
[00:27:16] Kate: Oh, it’s a good thing that you aren’t ever going to have to think about it, because as a community interest company it’s not a problem, people don’t throw money at us regardless of what potential for innovation and impact we have. If you are not delivering returns for investors you are just not accessing the same funds.
And that, I would argue, is a problem. But that’s the reality of the world we live in, and we’re not naive about that. So, anyway, that was just a bit of a.. I was just curious to see what you would say.
[00:27:50] David: I mean, the intellectual, mathematical challenge of this would be wonderful to actually dig into, but let’s keep going on these discussions.
[00:27:58] Kate: I know, so interesting. Okay, thanks, David.
[00:28:01] David: Thanks.

