AI safety, RAG benchmarking and responsible AI at ACM FAccT Conference

Watch the episode
A graphic with a grid background and a stylized flowchart in pink and blue.
Episode 6: AI safety, RAG benchmarking and responsible AI at ACM FAccT Conference

In Episode 6 of Mixture of Experts, host Tim Hwang is joined by Vagner Figueredo de Santana, Marina Danilesky and Shobhit Varshney. Today, what’s the future of AGI? The experts unpack Leopold Aschenbrenner’s AI safety screed, Situational Awareness. We also break down the state of responsible AI amid the annual ACM Fairness, Accountability and Transparency (FAccT) conference. Finally, we chat RAG benchmarking and what it tells us about the industry as a whole.

Key takeaways:

  • 0:00 Intro
  • 1:48 ACM FAccT Conference
  • 15:45 AI safety and general AI
  • 29:26 RAG benchmarking

The opinions expressed in this podcast are solely those of the participants and do not necessarily reflect the views of IBM or any other organization or entity.

📩 Sign up for a monthly newsletter for AI updates from IBM.

Episode transcript

Tim Hwang: I’ve never seen um like AGI being more plausible than we are standing right now so my hot take where we are right now 2024 um 5 years out I would see us um be able to get to very very intelligent machines hello and happy Friday you’re listening to mixture of experts I’m your host Tim Hwang back again each week mixture of experts distills down the week’s most important headlines and chatter in the world of artificial intelligence from research papers and product announcements to ethics governance and just plain gossip we’ve got you covered this week on the show first the annual ACM conference on fairness accountability and transparency or fact is happening this week in Rio we’ll talk about the latest developments in ml fairness and the state of responsible AI next up Leopold Ashen Brenner’s AI safety screed situational awareness hit the airwaves with a widely talked about inter with dorish Patel what’s the best way to forecast AI capabilities and what’s going on with safety at open Ai and finally benchmarking benchmarking benchmarking this week we talk about the latest in rag benchmarking and what it tells us about the industry as a whole as always I’m joined by an incredible group of experts who will help us cut through the noise and drop some hot takes as we go Vagner Santana staff research scientist Master inventor and importantly debuting for the first time on Vagner welcome to the show.

Vagner Santana: Thanks for having me.

Tim Hwang: Uh next up Marina Danilesky research scientist welcome back to the show.

Marina Danilesky: Thanks happy to be here.

Tim Hwang: And Shobhit Varshney who has been with us since episode number one uh senior partner Consulting on AI for US Canada and Latin America Shobhit welcome back to the show.

Shobhit Varshney: Absolutely love these thanks for having me again.

Tim Hwang: All right well let’s just jump right into it so the first story I want to cover is the annual fact conference is happening this year in Rio for those who don’t know it is uh arguably the leading conference on topics of machine learning fairness and responsible Ai and I thought this would be a good jumping off point just because if you’ve been watching the space for some time responsible Ai and ml fairness has become kind of a buzzword that lots and lots of people have used in recent years and I think these conference is a good time to check in on what the kind of state of play is In fairness and accountability questions in Ai and Vagner one of the reasons I want to have you on the show was um you’ve been watching kind of the papers and the chatter around the conference maybe I can just kind of toss it to you first for our listeners uh any sort of patterns or trends that you’ve noticed I think this year um fact if there’s particular papers that you think people should check out just curious about your review or your kind of thoughts on um what you’re seeing out there um uh at this year’s conference.

Vagner Santana: Sure uh well there there are interesting uh discussions around um uh synthetic data around how uh uh how people are using llms to create data and then uh also to assess llms using llms so they’re there’s this discussion going on also about uh responsible AI uh one of the papers I I I selected to discuss with you all I think has to do with how to how people are learning about responsible AI on the job I think that that’s important because people are getting interested and people are uh following and trying to find uh resources but then that comes with all the the uh the complexities of working organization so that that is one aspect as well and well the other aspect connects with a copyright and also how to deal with uh uh all the labor that is being packed by uh the the the use of LMS in a wide range of of uh jobs around the world.

Tim Hwang: Yeah for sure and I did want to pick up on that second theme specifically you know there’s much more as through with all these conferences there’s many more papers than you’d ever have time to discuss um but I think what’s so interesting about the responsible AI topic is you know this is really an evolution I think In fairness and ml where I would say even a few years ago basically a lot of the attention was like can we define in computational terms what a fair machine learning algorithm is and it kind of feels like there’s a lot more work now that’s happening in this much broader question which is okay well we have all these techniques and approaches around fairness and ml how do we actually get like an organization to implement it how do we get people to learn about it what are the techniques that people use and you know Vagner I think in addition to your research it kind of sounds like you’ve been you’ve been doing some work on this internally within IBM and so kind of just curious if you want to talk a little bit about that paper that you mentioned and then just kind of map it to your own experience I think I’m curious about like what you’re sort of learning as someone who’s you know very much in the trenches you know trying to get this work to work.

Vagner Santana: Sure yeah and one one of the aspects that um the paper curves and and has to do with incentives and we need to be aware of the incentives that our organizations have before uh thinking about responsible AI because otherwise U well we’ll be facing a lot of blockers along the way um and also uh there’s an interesting aspect that the paper highlights about um the the discipline identities that we have when we are like in in hard technical teams they have their own discipline identities and uh when they’re looking for uh let’s say resources about responsible AI they’ll probably go into resources they are used to look for so they they’re going to be looking for Technic technical libraries or metrics and sometimes we need to go beyond this own discipline identities in and look for other skills and other uh disciplines and to learn more about let’s say social technical impacts right and Beyond going Beyond focusing on let’s say some faom matrics for folks uh focusing on more technical aspects and the other way as well right for people uh of thinking about uh let’s say indirect impact on society they also need to be aware of the daily job of data scientists and coders and Searchers and how can we like do this a connection right.

Tim Hwang: Yeah I think your just rundown I think runs into or I think highlights I think a bunch of the issues uh you know I think it was a it was a joke that I had for with a friend for a while that like oh the main thing about a lot of big companies would do when they wanted to do like ethical AI or responsible AI would be like well we’re going to create like this like secret group of lawyers that will just determine everything um and it was like this is like not a good way of of doing you things um and I guess I’m kind of curious I mean Marine if I can bring you into this discussion you know um I guess maybe one thing I’d be curious about is like how you all think about things like fairness and responsible AI in in rag right which you’re literally trying to pull information from another source um and and I guess I’m kind curious about like if you’ve kind of you know have thoughts on this particular discussion right like how should organizations kind of best organize themselves to to do this right because I think part of it is this kind of interdisciplinary cross talk which I think organizations of different size you know do better or worse at you know in different capacities.

Marina Danilesky: I think uh something that Vagner said about incentives really pops up here as well which is why should you care well it’s because you would like your customers finally to be using your rag system and if it is giving answers that are not um not even so much fair but there’s a risk that uh it’s going to give something that is irresponsible that is going to lead to your users being misled or being upset or taking legal action then your solution is not going to be bought it’s not going to be taken so actually because we usually are looking at Enterprise use cases we are very very incentivized to make sure that we are communicating things that are you know Fair ethical regardless of what our own ideas are it’s because if we don’t succeed in that it will not be purchased the risk is too high um there’s too many you know fun stories in the news about uh what happens when you don’t pay enough attention to that.

Tim Hwang: Yeah so you’re actually seeing that because I think I don’t know I had a fear you know um which I still kind of have which is like maybe this discussion is going to become a little bit like um like data privacy where like I think early on there was kind of this idea that like oh well the minute there’s a really big data breach then everybody’s suddenly going to care about data privacy and security and like consumers will all prefer the the better privacy option right but then I think you could make the argument that one of the things that’s happened is that there’s just like so many huge data breaches now so many big failures that like almost like the Overton window has shifted we’re just kind of like oh you know someone leaked billions of customer records I guess that just happens but that is something that you’re seeing it kind of sounds like that like at least In fairness because we’ve had all these high-profile failures it’s not necessarily people have just become resigned to it it actually Still Remains kind of like a thing that people are really concerned about.

Shobhit Varshney: Yeah we seen this quite a bit right if you look at uh the AI culture the AI framework responsible AI framework on what would you prioritize and which ones are high risk how do you categorize use cases and so forth and there’s actual tooling and platforms that are needed to go drive these at Enterprise right and those three layers have to be addressed one by one the AI culture around hey look at your day-to-day workflows and see where you can apply Ai and you have to do this in a responsible way and here’s a framework around it unfortunately the reality on the ground for most of the Fortune 100 companies that I work with the responsible AI team you have to go it’s easy to go create a governance board and you go to the governance board for guidance and coaching and making sure that you’re not doing the wrong things unfortunately it becomes um I’m going to quote Lord of the Rings g off standing on a bridge and saying you shall not pass right go back to the shadow so we’ve we’ve somehow created a a forcing function that anything that goes to the governance board adds about 2 months of delay to a project so the value the unlock for the business gets dimin like hey I might as well not when deal with this and I should just go stick to my rpas or automation scripts or regular AI stuff and that’ll be just fine right so they’ve become a rate limiting step at this point and that has to fundamentally change and for that the next layer that I was talking about in terms of platforms that becomes more and more critical so instead of saying that hey you need to go figure out all these 20 different checklists in your rag pattern so I can know exactly where the data is coming from there’s uh there’s metrics that I need to report against and so on so forth now you start to move towards a platform approach where you say use the platform pre-approved accelerators all the rag when we looked at rag patterns within IBM Consulting within a week’s time we had like 121 different different ways in which people were doing RS and we said guys time out we got to go consolidate we’ll create scrip flow we’ll create a mechanism that has the best of all techniques in one single spot right so when you start to get to a platform then you come to a point where the governance boards are pointing you towards accelerators versus becoming a you shall not pass moment right I think that that whole culture leading to governance then leading to the stack and we’re doing this with a lot of our Fortune 100 companies recently last a couple of weeks back I had Pepsi on stage with us where we talking about how we’re helping them build a culture of responsible Ai and the Frameworks and so on so forth this is one of the many examples where we’ve had to go do this end to end from culture to Frameworks to actual tooling that goes and deploys that.

Tim Hwang: Yeah that’s really interesting and actually I should add that I’m surprised that it has taken this long for us to get to a Lord of the Rings reference so I think we’re episode six right now this is the first one that we’ve actually uh heard yeah exactly well and I think I don’t know I mean maybe one last nuance to kind of touch on I’d be curious to get the panel’s thoughts and V May back to you is like you know so for the last few episodes we’ve all been very excited about open source and it feels like part of the problem of Open Source is that you know suddenly like your fairness methodologies are almost competing with like just being able to like pull something off the shelf and like deploy it in any recess way that you really want to um and I guess I’m kind of curious about like how we think about sort of responsible AI going forwards in a world where like anyone can just pull AI off the shelf and use it um because it feels like in a world where like maybe there’s only two or three platforms you really can say okay well if you want access to this advanced technology you’re going to have to go through this additional compliance cost even if it takes you a little bit more time but that kind of lever is I don’t know from my point of view seems to be like breaking down a little bit as it becomes more and more accessible I don’t know if you buy that you might also just say Tim you’re totally wrong but I don’t know vager if you’ve got any thoughts on that or anyone really.

Vagner Santana: Yeah for in terms of uh open source I think that an interesting aspect is that uh people um have more transparency as as we I know and and also thinking about uh well fully open source models because people are also discussing that when you also only have the the model and you don’t know the data used to train the model you just have have have open source model and when you know more about the model then you have a fully open source right and I think that that’s important and people are getting more and more interested on on that and when we talk to clients they are also interested on um finding the right model for the right task I think that that is also interesting uh that pattern that is emerging like people discussing okay is this the right size of model for solving my problem is is this uh like is this language model or generative AI fully open uh can I host that in my own private Cloud so these are questions that are appearing when we talk about responsib AI right now.

Shobhit Varshney: Yeah we again I’m coming in from a very Enterprise approach to this like my my Square focus is how how do we scale these right and you look at a step-by-step process any workflow that’s happening in an organization today right seven different steps uh step number one you’re going to pull some data step number three you’re going to do some fraud detection step number four now you’re going to go extract something from a document invoice a contract something that came to you right now say we I was able to do OCR and pull that out with about 80% accuracy so far that’s best where where we were now all of a sudden we have llms and we say hey I have reason to believe that I could potentially get about 90 92% accuracy I can squeeze more out of this document right so now you’re saying about 10 12 points of of additional benefit that you can you can derive from it right at that point we stop and say if you have reason to believe an LM could do this let’s talk about constraints the constraints around cost envelop right how much can I afford if I’m doing this a thousand times or a million times there’s different Roi attached to it right then there is security of where the data resides models follow the data data gravity they go we deploy them closer to where the the restricted data sets are and so so forth then you start look at how quickly you need an answer the of a model matters you talk trying to start figuring out uh from a compliance perspective will I have to go explain to somebody how I came up with this answer which means I need auditable responses I need more deterministic responses in certain use cases and so on so forth so you come up with a set of constraints and given those 5 10 different constraints now you have two or three good athletes that you start to test with and then from there on we start to move towards metrics and see which one is is giving me more versus the others and so on so forth but it’s very critical at a step level at a subtask level LEL you’re trying to figure out which LM is going to do the job and we’re getting away from hey earlier we said hey can I can a gp4 model do the entire work for n to end so we’ll talk about that in a little bit but I think we’re still at the subtask level we’re surgically infusing Ai and geni and seeing if it can do this one thing incredibly well I’ll take care of the rest before and after.

Tim Hwang: Yeah I know I think it makes a lot of sense and I think goes to this really interesting question which we won’t have time to address today but we should do on a f future episode is you know what’s that mean for responsible AI right because it’s like you know you have lots and lots of subm modules that may have you know various various different types of problems there’s almost kind of a question about whether or not any one deployment is responsible but then whether or not the whole system hangs together is a whole another set of analysis right that actually is is another question great so I want to move us to the second topic uh of today um this is a really big week if you track the discourse around artificial general intelligence Leopold Ashen brener who is a former open AI super alignment team member published this massive online screed called situational awareness um this would have kind of existed I think as sort of a weird uh obscure screed uh but Leopold ended up doing an interview with dwosh Patel the sort of influential Tech podcaster and this story and this document has now just gone everywhere so I’ve caught up with friends who you know work in policy in DC saying we’re getting calls from congressional offices saying what our take is on situational awareness so I just want to take a quick breather here um uh because the claims of situational awareness are quite breathless um the argument is if you take all the existing Trends in Ai and you project linearly we will reach a point where AI becomes you know sort of um uh transformational in its impact and so I think this is kind of a great opportunity to bring in you show bit to this conversation because the way I sort of see the discourse is that there’s a circle of people who are in like AGI super intelligence land right who are like the AI is going to take over the world right but then I think like there’s this vast group of other people who are just like doing work with the technology who are like talking to companies that are implementing the technology and I’m kind of curious as someone who’s like really right you know at the front lines of that like our companies like situational awareness do we need to be worried about AGI like does that even enter into the commercial discussion or is this like a completely like almost in the parallel Dimension.

Shobhit Varshney: I’ve never seen um like AGI being more plausible than we are standing right now so my H take where we are right now 2024 um 5 years out I would see us be able to get to very very intelligent machines now the definition of AGI has been very wake right everybody has their own interpretation of what artificial general intelligence would mean and even if you compare two different people it’s very difficult for us to really have a good metric on is this person show with really intelligent or not right like if you ask my wife or my kids you have very different answer so it’s we’re very different point and even being able to Define what AGI looks like right but if you just talk about intelligence uh we’ve been doing a incredibly good job at making progress every two years if you like stepping away from the half a year increments is looking at a two-year horizon right when uh gp4 stopped training and you know they they’ve discussed this in 2022 you’re looking at about a half a billion dollar spend about 10 megawatt uh what 25,8 a100 gpus from Nvidia at that point right that’s kind of what they must have spent doing this 2024 today you have you can have 100,000 h100 equivalence you’re seeing what how much investment meta and others are making into this right and then you you had this huge big announcement with open Ai and Microsoft that they going to establish a hundred billion super computer right now we’re talking about something that starts to get into 2026 time frame and you can potentially have a gigawatt cluster you can have this big big giant machine and the power needed for it would be equivalent to say the H Dam right so or nuclear reactor right so now you’re trying to start to say that I can solve a lot of trough Problems by throwing more compute at it that’s just part of the equation right there’s better algorithm algorithms there better data that’s needed It’s a combination of those you can’t overcorrect for bad quality data with having more compute so we’re getting to a point where now you would get to more and more compute power being available if you keep extrapolating that out I see a situation where we would have have more than a nuclear reactor attached to this one of these big machines and you can then have a huge cluster that just intelligently look crunching through numbers I think what he’s extrapolating was by 2030 will be at 100 gaw I think that’s a stretch that’s about 20% of US Electric production but it does uh bring in in a few different aspects the safety of the AI who should have access to it Nations versus private sector we solved for that with the nuclear energy saying hey only the the big National governments should have access to to nuclear power right we to nuclear Arsenal and then we trust that there’s a mechanism in place that has checks and balances in the government that has access to something that’s super foundational such a massive impact in humanity it should be in the hands of of governments but you start to look at some of the world leaders around right now a lot of them and potential elections coming up and stuff too they don’t quite understand what we’re dealing with I’m just if looking at the access of AI I would I could easily see a geopolitical issue here where the country that has has those clusters if you think about the opheim if you go back in time and look at what we did during that stage you would not want to have that entire establishment in a different country right us went out out of our way to ensure that that’s being built inside the United States right so you’ll see a lot more of concentration of AI superpowers and how much they’re investing in building the energy the requirements building these massive clusters and if you follow the trajectory of electric production I was um I was giving a talk recently on how much was the impact of on sustainability perspective of of AGI and super computers and stuff and I had looked at this this detail around the per capita electric generation right how much elect does each of the countries generate and if you just look back at the last 30 years us United States has declined 5% in electricity production per capita United Kingdoms the UK has gone down 20 3% in the last 30 years China has gone up nine times in in energy production per per capita right so you’re starting to see axes of power that who has access to what kind of energy who has access to what kind of compute power and then to your earlier Point Wagner once you start to open up these models and open weights being available you’re essentially giving people people the recipes of how you can go replicate these things on your own right so I think we’re we’re at this this weird intersection of private versus government and then do does the AI intelligence then dictate geopolitical power and when does that tip over at what point does the government start getting really really serious about safety who has access to these Technologies inside of open AI or the big Tech Giants and things of that nature how open are you about what data is going in and so on so forth I’m I’m just very fascinated by the impact it’s going to have.

Tim Hwang: Yeah it’s actually I don’t know I feel like you you you surprised me there actually right I thought you were going to go in a completely different direction I feel like when I talk to many kind of folks F who are like in Enterprise on the business side of this they’ll basically say this is not happening this is not realistic like you see what’s happening with AI right now it’s never going to be like what this guy leapold says um and and it feels like you’re actually going the opposite direction you kind of say look you take all the existing Trends you extrapolate them out and we’re going to be in a really weird place in 24 months I guess Marina Vagner I’m curious if you to sort of like agree with this kind of assessment or you know from the researcher side is it right to say hey these linear Trends are basically what we should use to think about capabilities in say 2026.

Marina Danilesky: Tim we all know that nothing wrong has ever happened from linear extrapolation in the history of humans very dependable very dependable this is always how things go um I I think I do have a bit of a different perspective than chit and maybe a little bit more like the one that you had said where yeah I I don’t agree with all of the linear extrapolations and of course we all can have the the perspectives that we have on how things are going to go but I think that even if you continue to throw more compute more data the way that AI is currently implemented and we’re just in another wave we’ve gone through waves before we’re in the current wave there are to my mind limitations to what you’re going to be able to achieve and it is not completely clear how you will actually get out of an AI never recommending you to put glue on Pizza just because you gave it more computer and more data um and so I think that while we are closer we’re still not there and in my mind there’s at least another one or several technological waves that need to come before we really get there so is there going to be a lot of interesting things coming sure the points that shet raises about uh accessibility and who gets to actually have these models that has a lot of really interesting uh implications for being able to disseminate misinformation have an impact on how people perceive information and so on so forth do I think that that’s going to get to AGI personally no but it doesn’t mean that it’s not going to get to uh places that are very impactful.

Tim Hwang: Yeah and I think there’s actually one thread in male throat to V I’m curious about your thoughts that I hadn’t hadn’t really thought about which is very very interesting is kind of I guess this is kind of bet about like what does compute actually get you right um there’s kind of one view which is if so long as I feed in more data and more compute the representation in the model will eventually just become accurate like we’ll solve the just eat rocks by basically like you know kind of like Computing our way out of the problem I guess Mar you’re kind of saying I don’t know want to mischaracterize you that sort of like there’s actually some genuine questions as to whether or not that that will even happen right like the models will become more powerful but they might not necessarily become more accurate right we normally think about things getting better as you know kind of trending in a certain direction I guess you’re saying we can see Improvement but it might be very multi-dimensional in a way that kind of is a little bit counterintuitive I think.

Vagner Santana: Yeah yeah I think that the the the issue with uh increasing more or requiring more compute to improve or increase uh the already really large models uh we’ll we’ll be seeing like less and less uh organizations controlling everything so that’s in terms of responsibity I think it’s it’s something that may be concerning and in terms of environmental impacts also people are thinking a lot about the energy that uh these models are not only a requiring for training but also inference at scale right so then these I think that balancing all of these I think that that’s a big Challenge and in terms of responsibility right what we are always trying to think about is uh do we really need to create this technology right now what are the problems that we’re going to solve uh can we so solve the problems that we have with the technology we already have right there’s a lot of interesting questions and uh uh well in terms of responsibility we need to think about these all the time not only as after T right that like part of the responsibility might just be like no AI I think the cost and impact of this is going to start plementing right.

Shobhit Varshney: If you just look at the compute power that’s in your phone uh today uh that costes plementing so over time we’ll solve for this I think from Enterprise perspective Tim your original question I think we are over complicating how work gets done in organization if you have access to hundreds or millions of MIT Harvard and Stanford grads and you put them into something very mundane and say you’re going to do procurement analysis you’re going to get an invoice you’re going to compare it against something right that’s the kind of work that happens in an Enterprise right so I have reason to believe that if you put a really intelligent person or an equivalent of a person digital labor inside of a particular workflow that subtask will get done very very well the all kinds of guard rails and stuff that you can create around that particular task so if you look at in levels the first level is can I do a subtask really well in the previous discussion I said step number four I’m going to extract something out can I do that toss really well and that starts to become a specific unit of work then you go one level up and say today a human asks each Mach each llm to go do different steps can I replace that with an orchestration where an llm agent can figure out a plan manage the memory and stuff and automate the entire flow into to end there’s a very plausible path for us to get to figuring out how step is done Auto orchestrating all of those workflows and now you start to move up the hierarchy of what a human supervisor would have done versus a summer intern would have have told you’ll always double check what a summer inter does that’s where we are today and over time you see a progression towards work itself getting automated to with a very very high accuracy especially with the cost of AI plementing over time.

Tim Hwang: Yeah that’s that’s almost a great way of thinking about it is to show your earlier point about sort of AGI having like this very amorphous definition uh it’s almost interesting thinking about the idea like William Gibson as his quote right like the future is here it’s just not widely distributed yet kind of what you’re saying is like AGI is here it’s just not widely distributed yet like for certain types of tasks like the the AI that we have right now can do all the possible job tasks right um and you’re basically just kind of talking about like how far up the organizational chain this thing will go.

Shobhit Varshney: I don’t think that I don’t think we’ll get into a point where we’ll have a crisp definition of what AGI is and we’ll say hey today ra open a champagne we reach that right it’s incremental progress and a different definition in each field in each domain in each task right so I think the right the right frame to say that machines will get super intelligent over time and they will exceed human intelligence in certain tasks and one thing that humans don’t do really well is share our knowledge amongst ourselves right if you put two experts in a room it’s very difficult for them to actually go at a problem together right we don’t do a really good job that’s expanding out using the network effect I think that’s going to change when you have super intelligent machines that can talk to each other and and and drive better safety better algorithms better research and start to build better algorithms all together right so I think I’m very excited about the direction that we’re going so I want to move us on to the final topic um uh and the way I want to te this up is that there’s this famous uh clip of Steve Balmer when he was CEO of Microsoft where he’s like if you’ve seen Steve Balmer before he’s like this big muscular guy he’s like very sweaty on stage and he’s just shouting developers developers developers and.

Tim Hwang: I kind of feel like if you had played that scene again today people would be being like benchmarking benchmarking benchmarking um because I think it is becoming such an important aspect of sort of like the supply chain of AI um and you know there’s lots and lots of things we could talk about with benchmarking we have and we will continue to um but I think Marina particularly with you on the line I figured it would be great to kind of zoom in specifically to rag um and do a little bit where we kind of talk to the listeners about essentially what’s happening in rag benchmarking um and then I think from there kind of talk a little bit about what that tells us about how benchmarking in the industry is evolving uh not just in the industry but I would say in research as a whole but um if you if you will I wanted to kind of throw it to you and to say if it’s possible we’d love kind of a short crash course into how people think about measuring the quality of rag and that will almost give us something very concrete to talk about in terms of benchmarking generally.

Marina Danilesky: Sure sounds good so I will say benchmarking that’s been around for a very long time it’s always been something that was extremely important for systems databases ml everything so I understand folks maybe looking at that right now but yeah you’re you’re into it before it was cool that’s right that’s right we were into it before it was cool we discovered the band first um so the thing right now with uh rag let’s talk about what rag is again real quickly and then we’ll see what it is that you need to evaluate the retrieval the augment to the generation part all right so what are you trying to do you’re trying to finally give information that is supported by uh knowledge that you can say okay this is knowledge that I can assume here’s information I’m giving you so what happens with rag remember uh user has some sort of an inquiry you fetch something that is related and you say I’m going to give use this information to give it the answer okay where is all of this going to break it’s going to break when the query is not well formed so you’re not fetching the right thing if you’re not fetching the right thing then you don’t know that there’s information you didn’t get at so that’s something to evaluate if you are fetching the right thing or even not the right thing you then have to generate an answer based on that multiple ways that that’s going to break down so you’re going to have a model that gives an answer that’s not based on the information you fetched it’s going to give you an incomplete answer it’s going to give you an answer that’s a mix of some of it is drawing from it some of it is drawing from its parameter some of it is just making up because it decided to go off uh especially later in the response and you have to check all of that uh can you force the model to give you a different answer because you told it no no no you told me it was this way but I’m gonna say now assume it’s that way okay can you can you mess it up that way can you uh give the answer quickly enough um so these are the things that you have to manage to evaluate so when people talk about context relevance or um answerability or the faithfulness or completeness all of these different metrics that people have uh this is really what we’re talking about evaluating rag couple of points here you can try to Benchmark uh with a against a gold answer which is usually something that works in cases like classification or anything where there is a very very clear thing as an answer the problem with generative AI is remember that word generative everything is created fresh which means that there might have been a lot of different ways to create an answer so when you’re saying that I’m going to have some sort of an overlap metric like Rouge or blue or anything of that kind that’s not always going to be great it’ll tell you if you’ve gone off completely but it won’t tell you subtleties that oh maybe you rephrase the answer a little differently but it still would have been acceptable so problems there so then you say okay let’s not have references let’s just judge the answer as it is the problem with all of the metrics that I just mentioned nothing has a compl completely clear definition and it can’t because you cannot get everybody to agree on what does complete mean what does faithful means believe me I’ve tried we have had so many arguments with research well I mean it’s kind of essential because like yeah sorry go ahead no it is you’re completely right it is like what does it mean for you that an answer is complete not only can the researchers not agree then the customers can’t agree so when you are talking about benchmarking uh there are bits that you try to Benchmark first parts of this system as shet was saying well how do you do on just the retriever part how do you do on just a generative part how do you do on you know just faithfulness and the problem is that here the whole is not the sum of its parts you put all of that together in an end-to-end experience and it is not equivalent to I checked every part individually therefore I know how it’s going to go together doesn’t go that way and it’s a very difficult thing to actually Benchmark because the more parts there are to a system the more complex it is to know what happens when you put all of them together in different ways so that’s actually why people are so interested in the benchmarks right now is because the state of it is a little confusing it’s a little bit incomplete where just like what is it that we can actually trust and then of course what we talked about in previous episodes that the benchmarks do get saturated very quickly soon as you have one out a few months later okay everybody already can deal with that one you know you got to thank it for its service and and move on.

Tim Hwang: Yeah what I love about this is that it’s like it starts very tactical and then becomes like existential very quickly where you’re basically like what is truth what is Clarity anyways you know of which like there kind of is no answer I guess so know if Marina this is a good way to sum it up I mean are you sort of saying that there is no rag Benchmark in a certain sense right like that there’s no commonly understood Norm for judging rag quality.

Marina Danilesky: We do we do our best and I think there are incremental uh implementations that are better and better and better as we have one Benchmark realized something it didn’t cover do another one do another one do another one so there are you know in incremental approximations of what is and is not going to work and at some point in time again it’s probably going to reach a level where we say all right this is good enough we’ve we’ve kind of you know saturated this as much as uh we can but what ends up happening is then you end up you moving to other use cases right shet mentioned agents it’s a very interesting direction that we’re going in well now you don’t just have text you don’t just have that at going out into the rag now you have I am calling functions I am using tools I’m having something else happen in the middle my execution plan as an LM agent is absolutely all over the place now you don’t just have an r and a g now you have I don’t know how many things every sing single time you add now how do you Benchmark now how do you Benchmark so we all are having a lot of fun constantly making new problems for ourselves that we then have to test that then reveal additional problems and and things we can Implement.

Tim Hwang: Yeah I love the idea that kind of like um eval design itself is trying to Hill Climb like basically like yeah has like a very similar pattern to the evals themselves so um yeah.

Shobhit Varshney: So Tim just working with real clients one of my big big clients we’re looking at contracts and rag is a great example of that right given a whole bunch of thousand a few thousand contracts and would ask questions against it expect to get a good answers right so when you start to look at the kind of questions and queries people are going what’s going to be insightful for them there’s a a level one question is can I find something in a contract that tells you what’s the expiry date of the contract or is is there an exit clause in this contract or not that’s a simple rag pattern right very naive and can work then you start to look at this is a contract but then it has amendments stapled to it and now the answer of the end date actually is in the third amendment that overrides the previous right so now you’re looking at the whole Chain of Thought of how to read this particular document then a level three of of a question could be when I’m trying to cross compare and say hey I want to order another thousand units which one of these uh contracts is closest to the threshold where I’m going to get some cash back it’s a more complex question and just very quickly start to move away from a rag so the perception is that oh I can ask I can dump contracts ask questions but in reality a human would have gone and looked at another system in an sap and said here all the orders to date and then that’s some math on it and then giving you an answer that’s going across so it’s not quite right there’s no document that gives you the answer that you can go retrieve on demand so you need to have some type of a router in the middle that understands what kind of question is asked and then you may have to go chat with some structured data at the back end to bring that in and then call unstructured starts to get really complex we talk about rag but we should really be talking more about the use case and to end that has much more than just the rag patterns need be metric so that’s more complex than what Marina you were talking about now we’re talking about the whole end to end chain and how how do you measure accuracy in this case two different people have two different answers.

Marina Danilesky: Yeah see we even disagree because I call all of that rag in my mind and so like we don’t even agree what rag patterns are because to me I’m like great you’re you’re retrieving a function answer you’re retrieving something from a knowledge base you’re you’re still kind of you know retrieving this Cas just means you know function call but so even with that right you can think of rag pattern as just a single call and only informational cleer you can think of it as the entire thing that you’re talking about shet and I think that you end up having to extend how you do the evaluation great we’ve done it for one small pattern now how about when you extend extend extend extend so yeah you’re right Tim hill climbing what why why sit on our Laurels when there are more complicated problems evals we could be building yeah.

Tim Hwang: And I mean you know my bias is just like I think one of the things I’m most interested to see in the AI space is just like the continued growth of evals as an industry because this is like where the endless value will will emerge right where like company’s being like is it good and it like actually ends up being like is very very deep question that really requires some real sort of craft and expertise um so Vagner you get the privilege of having the last word on the episode as our uh inaugural or sorry our debut guest uh this uh this episode um any final thoughts uh on kind of the benchmarking question or rag in general I mean I’m always excited about rag hot takes um.

Vagner Santana: Well now not a word that relates to to rag just a final word that I like to to uh uh more a provocation like so for folks interested on on responsible AI I think it’s worth to try to go beyond your discipline identity your bubble of content and try to reach out to other contents because and we’re talking about fact and fact is interesting because they go to a more Technical and also more to the humanity so try to find a a subject that you’re interested like rag or other uh uh subjects and try to go Outsider discipline identity I think it’s it’s good for for for uh um theom as a whole.

Tim Hwang: Yeah for sure that’s a great note to end on well uh that’s all the time we have for today uh Marina Shobhit thanks for joining us again.

Marina Danilesky: Thank you so much for having us Tim this is awesome most fun thing we do every week.

Shobhit Varshney: Yeah definitely thanks for joining.

Tim Hwang: And um Vagner um thanks for joining and hopefully we’ll have you back again sometime.

Vagner Santana: Thank you.

Tim Hwang: Great well if you enjoyed what you heard you can get it on Apple podcast Spotify and uh good podcast platforms everywhere uh and we’ll see you next week.

Stay on top of AI news with our experts

Follow us on Apple Podcasts and Spotify.

  1. Subscribe to our playlist on YouTube