275 – The Agency Fund’s Social Sector Tech Stack: The Back End

The IDEMS Podcast
The IDEMS Podcast
275 – The Agency Fund's Social Sector Tech Stack: The Back End
Loading
/

Description

David and Santiago discuss The Agency Fund’s recent article on a technology stack for the social sector, exploring the backend systems that make modern AI-enabled tools possible. From LLM gateways and agent builders to data pipelines, monitoring, and experimentation, they examine the growing ecosystem of open tools that can help small organisations access capabilities once reserved for large technical teams. Along the way, they reflect on how these developments relate to IDEMS’ own work and why this is an exciting moment for organisations seeking to build impactful technology with limited resources.

https://theagencyfund.substack.com/p/a-default-tech-stack-for-the-social

[00:00:07] Santiago: Hi, and welcome to the IDEMS Podcast. I’m Santiago Borio, an Impact Activation Fellow, and I’m here with David Stern, one of the founding directors of IDEMS.

Hi, David.

[00:00:17] David: Hi, Santiago. Looking forward to the discussion today.

[00:00:21] Santiago: Yes. What’s it on? I don’t feel very well prepared for this.

[00:00:25] David: That’s okay. Just last week, The Agency Fund, on the 25th of June, 2026, released a blog post, an article, on “A Default Tech Stack for the Social Sector, The AI tools we recommend for nonprofits and why”. It’s an article by Shubham Sharma, and it’s very sensible and it’s very aligned with a lot of the things we are thinking about and we’re doing.

So I thought it would be sensible for us to just be in conversation around this, share it with others, maybe other people will hear about it, and also offer a few perspectives of how this relates a bit to our own work.

[00:01:11] Santiago: That sounds good. Where do you want to start? 

[00:01:15] David: Broadly, the way they’ve structured it in terms of the tech stack is they have the two underlying pieces, the AI stack and what they call the observability and learning stack, and then they have three front ends, as they put it, a frontline worker, mobile chatbot and custom applications. That’s broadly how they think of this or how they’ve organised this, which I think is very sensible.

So we should probably discuss very briefly in this episode, all five of those pieces and just relate to them, be in conversation with them, even if we are not gonna go into great detail.

[00:02:05] Santiago: Okay, the first two seem to be more like the behind the scenes engine, while the latter three are more about how we use systems. Am I understanding this correctly?

[00:02:21] David: Sort of, yes. The two backend pieces, the AI stack and the observability and learning stack, these are all about the underlying systems. This is where we now have these multi-agent systems, which are relatively easy to build using AI, large language models, small language models, other AI tools, as part of how we build things and we work and we interact and we learn from them in some sense.

And maybe that’s the place to start, I guess that is the place to start, and we can just dig into very briefly what they call the LLM gateway, because the whole point of this, why is this article come about, it’s come about now because everyone is trying to use artificial intelligence through LLMs to gain efficiencies in their work. And it is a really powerful set of tools which now exists and can be used even by smaller organisations.

Their claim is that even small social organisations can access these sorts of high powered technologies and that the intermediary tools are starting to be there now, which they can recommend, and produce this whole stack for you.

[00:03:48] Santiago: And that can reduce costs of these increasingly costly systems that are out there.

[00:03:54] David: Exactly. So the first thing they talk about is what they call the LLM Gateway. And the idea is that instead of building everything in, let’s say, Claude or OpenAI, ChatGPT, or Gemini, it’s really good, if you are small, to have a layer between that so that you could use any of those systems, and if their pricing changes, you can switch between them.

And so to be able to have something in between which you are building in, is really strategic as part of a long term strategy for a small organisation, not to be tied into one of the big players.

[00:04:33] Santiago: And let me see if I understand this correctly. What you’re talking about is a sort of orchestrator that can then orchestrate different models depending on what you’re doing and depending on what the current situation or the developing situation is with the different models that are out there.

[00:04:54] David: Yeah, it allows you to write code once and send it to different models. You can also use this to compare between models, to see which model gives you the better results, and so on. It’s the sort of thing that we try and do in our own work. We don’t necessarily use exactly the same things that they’re recommending, their default for this is something called OpenRouter, they also suggest LiteLLM. These are all very sensible, there’s lots of others which are possible, they have a whole list of other things that they look at.

And one of the things that they articulate very nicely is that they have what they call a default, they have an alternative, and when to deviate, and additional information. And they do this consistently for all of these different pieces of the puzzle. And the reason this is such a nicely formulated article is, if you are a small team who is not expert but engaged in this, well, this does help you to have this whole stack available.

And broadly we will have a few points at which we will offer some alternatives in this podcast, because I’ve gone through the list and I have a few that I would like to discuss. But, by and large, everything they’re saying is sensible, and really, I like what they’ve done, it’s a really good piece of work.

Under their AI stack, they have the LLM gateway, which is what we’ve already mentioned, which is this code layer which separates you from a given model so you can swap between models. But they also have an LLM monitoring and evaluation, a speech system, a speech evaluation, agent builders, they have vector databases, these are all part of that stack. And part of where the complication is, is that, really, to build these systems well you do need to have awareness of all these different things in case you need them.

If you’re going to build something where they’re speech to text, you need to know about the speech systems, and so on. And just to go through very quickly, the LLM monitoring and evaluation is something which can be really useful to debug and to identify when things go wrong. They talk here about Langfuse, Calibrate, Maxim as all being sensible options for this.

The speech systems, they’ve got a couple, they’ve got ElevenLabs or Sarvam. There are others that they talk about, and there are others that I know in the African context, which are coming up, which are quite interesting as well. But again, the idea is that a lot of the tools that social organisations want to build might be for people with low literacy, they might be where speech is really important. And there’s this idea of being able to then evaluate speech systems, and they have both Calibrate and Maxim as things that they mentioned already for LLM monitoring, but which are really tailored to speech evaluation.

And the final thing they have within this sort of AI stack piece is, well, so the final two things are, the agent builders and the vector database. I will just mention this very briefly because a lot of what we’ve been discussing in our team at the moment has been these multi-agent systems and how actually a lot of the places where we’re working now are needing multi-agent systems to become effective, and so we can actually build efficient systems.

And they have singled out both OpenRouter and Pydantic AI as two systems, which they would recommend by default. And then of course they’ve mentioned that Claude and OpenAI have these very good agentic systems as well if you are willing to get tied into a single system.

And we’re actually, as an organisation, we’re following both of these routes. We’ve got people looking at OpenRouter, who are building in that, if you want, independent way. And we’ve got other people who are just going through the Claude agents because there’s so much which just comes out of the box that you can do quickly and efficiently, and then you can rebuild it in a system like OpenRouter once you know what you actually want to build in a quick way, which you can do very well with the core two.

[00:09:13] Santiago: I myself am exploring OpenAI agents.

[00:09:16] David: Oh, you’re on the OpenAI agents, yes, so you are on the OpenAI agents, Michele is, I believe, on the Claude agent, and Lily and George are the ones using OpenRouter. And so we’ve got a whole. The idea is really that each of you now are gonna sort of learn about, not just learn about them, but use these different systems to then actually rebuild and regroup, so that we actually probably will consolidate around something like OpenRouter, but with the experience from these different systems.

[00:09:48] Santiago: And the experience means learning what works and what doesn’t work from the perhaps more mainstream systems like Claude and OpenAI.

[00:09:59] David: Yeah, yeah. And what’s really interesting of course is that this has been transformed in the last six months. Even at the end of last year when we were discussing this, I was just letting people experiment, we were saying that the systems are still in early days. And about February this year, particularly Claude had these huge jumps forwards with these multi-agent systems. And now it is very much the case that we need to be, within our team and beyond, understanding how we want to use these systems and how we can learn from, where we can evaluate between the different systems.

[00:10:40] Santiago: And we have several use cases for that.

[00:10:43] David: Absolutely. Yeah, there’s a number of use cases where we’re just having to get ahead of the line on this, so to speak.

And then of course there’s the vector database. I won’t dig into that, but what I will say is how you store the documents for AI is something where this depends so much on the system you’re just choosing to use. At the moment there isn’t something which I would, and this is part of what they articulate, this is something where building your document stores is just something which you can just do within the systems. So it is not something that many people spend a long time actually thinking about at this point in time.

[00:11:27] Santiago: Correct me if I’m wrong, but I think that this is what will potentially allow us to create much more specialised processes within different models so that we can perhaps start moving away from the all encompassing LLMs to something a bit more specifically trained for a particular task.

[00:11:50] David: That’s absolutely part of what the vector database is about. I suppose the key thing is that most people trying to use this right now, they’ll just go through whichever model provider they’re using and they’ll add to this. What we will be doing over the next six months to a year where we’ll be taking this a bit more seriously, is actually trying to think about how we build this across different systems in such a way that we can be really testing the systems much more.

And I suppose that really brings us to the observability in the learning stack, because a lot of this is about the fact that once you have AI systems working, it’s all about the learning, how is it doing, how can you improve it, what can you do, what’s the human input?

[00:12:34] Santiago: Can you catch errors?

[00:12:36] David: Can you catch the errors? Can you, not only catch the errors, can you help the system to recognise the errors it’s making and why so that you can get them out, put them into what goes in first, and those errors are then avoided in the future? So you are actually improving the system as you go, and it’s these sorts of learning pipelines.

And at the heart of this, of course, are data pipelines, and I do like the fact they’ve got Airbyte → BigQuery → dbt as what they suggest as a sensible data pipeline for not-for-profit scale data infrastructure. There’s so many different options here, but it is a sensible one that they’re putting forward. And I could dig into that because there are so many options there coming around, but what they’re suggesting is Airbyte → BigQuery → dbt, and it is absolutely sensible.

And it’s interesting that you do need that sequence. Again, this is where you think about the load which you need as a small organisation building these systems. There’s so many tools you’re needing to learn at the moment when in a few years, maybe in six months or a few years, there will be systems which would’ve put this together in a way which is so much more friendly for small organisations like social enterprises, like social impact organisations in general. I’m convinced of that, and if others aren’t going to do it we will, so it will exist in the next six months, a few years.

[00:14:05] Santiago: Yeah, I was gonna say, my understanding from previous discussions is that that’s partly what we are exploring developing, if nothing comes out earlier.

[00:14:15] David: Absolutely. Yes, exactly. There’s so many good open tools, it is that job of putting them all together into coherent wholes, which just really makes it seamless as Claude and OpenAI have done for a more commercial market.

The next thing, of course, on their list is dashboards. And I was really interested. They chose Looker Studio because of its nice integration with BigQuery, but their alternative there is Metabase, and they also mentioned Superset, which are tools which are so standard. I mean, we’ve been using these for years with the apps that we’re developing for parenting work, and it’s really interesting. We found that Metabase, when we started using the apps, we thought this is perfect, our partners will love this, they can have access to all the data.

And they can’t touch it, it really is not the dashboard, it doesn’t build the dashboards that they need. And so it’s really interesting, there’s a whole other layer which is needed beyond this as well to build further dashboards for partners, and that’s a whole ‘nother story. I guess what I’m saying is that, even with the stack that they’ve got to, and there’s huge numbers of different tools that they’re arguing are needed, it’s still not everything. It doesn’t quite get all the way in terms of our learnings working with pilots.

Then they’ve got workflow orchestration, this is what in our team, Ian does a lot of this. And they talk here about n8n, which is very sensible as a way to automate workflows, when something happens do something else, to automate these processes.

They then talk about data catalogues and they’ve got dbt docs. We, as an organisation, we’re still using Google services for this really. There are a lot of open alternatives and dbt docs is the one they suggest, but there could be many others.

[00:16:14] Santiago: And we considered switching to others in the past.

[00:16:17] David: We have, and we have switched to other things for certain partners, but nothing we’ve switched to and nothing we’ve used commercially or in the open ways is achieving what I think is needed here. And so it’s really interesting, this is a space, this is a place where I hope and I expect we will see further innovation in the next few years because of what will now become possible integrating with these AI systems. So the data catalogues is something which I hope there will actually be movement on more generally.

And I love the last thing they have or the last couple of things they have here. They have Ad-hoc analysis and Experimentation as the two last things here. And this is really interesting. The Experimentation, I’ll start with that, this is what we’ve been doing with Oxford University and our partners there, building in testing, which researchers can then use as part of this.

Now this is the industry standard in commercial systems to do A/B testing for market research. But, really, what’s becoming much more powerful and really important is to do this not just for market research, but for actual behavioural research and to turn this into academic studies, because actually this is how you can build evidence of what’s working and why and how, by having this simple experimentation A/B tests.

Now they have an open source platform from The Agency Fund, which I wasn’t aware of before. This is one of the things that was new to me, we’ve built this into our own systems, but they actually have what they call Evidential, and it’s something I’m gonna have to look into because this is something where we know the importance of this, we know why we’ve built it into both the chatbot work we’ve done and the app, the Open App Builder. But this is something where there is obviously this open source platform which now exists and which is trying to do this.

They also have small alternatives, Growthbook, and Posthog, which, again, I have to admit I didn’t know about, and it’s interesting I didn’t know about these because this is something we’ve been doing for five years with PLH. We’ve been embedding research and this experimentation into our work. And so it’s really interesting that these tools are now coming out as open tools, which are hopefully gonna facilitate that. So that’s definitely a learning point for me and something which I am looking forward to digging into and seeing how it relates to the things that we’ve built. 

[00:18:55] Santiago: And it is of course very important for us because we’re so evidence focused. Evidence based is one of our principles, and not just within our principles, but also in the collaborations that we have, and we want to have, research is a very important cog in the machine. We want to be able to evaluate the impact and the efficacy of the tools that we develop.

[00:19:25] David: Yeah. So I guess the thing which I find really interesting around that is that there’s a need to move the research agendas as much as there’s a need to move the technologies. Actually, this idea of embedding research in these technological innovations is something which is not yet as widely accepted, as widely published on as I would hope.

And so really this is something I’ve been pushing our colleagues at Oxford in particular, who are pushing the boundaries in other ways. And really, they’re the ones who I hope will help to get this sort of research really established academically. Because it is one of those frontiers that people have done work on, there is work happening but a lot of what’s happening is not getting the academic recognition because the methods aren’t well established yet. And so building the methods alongside these tools is, I think, really important, there’s a lot of growth potential there.

So let’s finish this part with the ad-hoc analysis piece, because what I love here is, I would argue, this is the only one where it’s drawn a blank. Their default is Colab notebooks. Now, if people know what that is, this is in the Google infrastructure, this is part of big tech. Yes, it’s free, but it’s the only area here where I would argue they haven’t found a single open source tool which does this.

And yet, this is exactly where a lot of our efforts have gone. This is where R-Instat for climate is playing this role. It is not yet playing this role here, but it could, with the right sort of development. And R-Instat will not be the right tool for reasons we could dig into. But it is exactly what this gap is identifying, that there isn’t the right tool for these ad-hoc analysis, which really enables people to go beyond your dashboard analysis, dig into the data, with specific types of data, and to do so with structure but with real open-ended questions where you can ask whatever you want of the data.

This is where your traditional stats package, this is what you normally need to do with that. People would be using R for this, they’d be using Python, but then there are tools which can make this easier to do ad-hoc analysis. Ad-hoc analyses aren’t that ad-hoc, there’s a standard set of things that need to be done. And so we do need to make these tools easier.

So this is one, looking at this, I was really interested that I see this as a gap, I see this as a need. It’s a need we already know a bit about how to fill, and I’d love to actually dig into this, but it’s an emerging field and this is broadly what they say, they have their default Colab notebooks and they’re saying emerging using Claude with dbt and Postgres MCP. But this is experimental at the moment, and we know a bit about how to do this better. So that’s one which actually motivated me to try and say, okay, if we do start building out this infrastructure, that’s a piece where we could really contribute.

[00:22:53] Santiago: Okay, that is a very comprehensive summary of all the backend. I wanted to ask you as I was listening to your analysis: a lot of this is AI related, but we’re thinking about embedding deterministic components to our processes as well. How does that fit into this model or this structure that they’re proposing?

[00:23:21] David: Really good question. And of course the beauty of it is it fits in really well because they’ve got the agentic component. So, in the agent builders that we talked about before, the key is then you decide what the agent builders are working on, what are they actually doing, what’s their role in it?

You are not asking them to do the whole, you are asking them to do components. And so building that around these deterministic pieces is absolutely sensible and feasible. And that’s exactly how we would suggest building these things up. And in many of the social impact cases this is what people would want anyway.

The people that you are working with who are often marginalised or in difficult situations or vulnerable in some way, so you don’t want to let a purely AI system loose on them in that way. You want elements of humans in the loop, you want elements of determinism in what is presented and what’s built in.

So this is exactly aligned with that because of the agentic approach, which is included in this. But it’s also why, if you look at this, this is overwhelming. We haven’t even got to the front end pieces. And my guess is that’ll be another episode now. But just this backend layer, who in a small social impact organisation with a staff of five people is gonna have the bandwidth to learn about all this? And is it right that it’s one person? This is the thing, this is why in big teams you have specialists who take on each of these components. You’d have whole teams just doing the speech to text component.

What’s incredible now is because of the way these agentic systems work, they’re strong enough now that you can, as an individual, take on this whole stack. I’m not saying it’s easy, I’m not saying that I’d recommend it for someone who doesn’t really want to get their teeth in, but you can. And I believe over the next six months or a year, this is going to be become easier and easier to do as we actually build out the systems, which then sort of help this stack to just work together, actually take what the agency fund has done, create implementations of this, which are maybe not quite turnkey, but almost, really getting it to the stage where you can have a system which is as easy to use as some of these commercial, the big players, but built on these open technology stacks.

So it’s really an exciting time, and really kudos and credit to The Agency Fund for producing this report and for moving the conversation forward through producing this report.

[00:26:10] Santiago: Right, on that note, I think we should leave it there and look at the front end in the next episode.

[00:26:17] David: Sounds good. I look forward to it.

[00:26:19] Santiago: Thanks.