Can Mistral make Europe a global AI contender? In episode 55 of Mixture of Experts, host Tim Hwang is joined by Chris Hay, Volkmar Uhlig and Kaoutar El Maghraoui to discuss the drop of Mistral Medium 3. Next, we analyze the AI chip sales that NVIDIA and AMD made to Saudi Arabia. Then, with IBM’s new ITBench and OpenAI’s HealthBench, we dive deeper into benchmarks for AI evaluation. Tune in to this week’s Mixture of Experts for more.
Key takeaways:
The opinions expressed in this podcast are solely those of the participants and do not necessarily reflect the views of IBM or any other organization or entity.
Tim Hwang: Mistral is France’s national champion in AI. Will it make Europe a global contender for the technology in the years to come? Chris Hay is a Distinguished Engineer and CTO of Customer Transformation. Chris, welcome back to the show. What do you think?
Chris Hay: The US is just falling in Europe’s footsteps.
Tim Hwang: All right, great. Volkmar Uhlig is Vice President, AI Infrastructure Portfolio Lead. Volkmar, what do you think?
Volkmar Uhlig: I think the judgment is still out, but I hope for the best.
Tim Hwang: Alright, great. And Kaoutar El Maghraoui is a Principal Research Scientist and Manager for Hybrid Cloud Platform. Kaoutar, welcome back. What’s your take?
Kaoutar El Maghraoui: I think Europe may not win the race to build the biggest model, like what’s happening in the US and China, but it has a big opportunity to still define the rules of the road.
Tim Hwang: All that and more on today’s Mixture of Experts. I am Tim Hwang, and welcome to Mixture of Experts. Each week, MoE brings together a world-class team of researchers, engineers, and product leaders to discuss and debate the biggest news in artificial intelligence. As always, we have a ton to talk about. We’re gonna talk about a big shipment of chips to Saudi Arabia, some new releases in the world of benchmarking, and a new experiment around AI-generated ads. But first, I really wanted to talk a little about Mistral Medium 3, which was a launch that happened just a few weeks back. It’s part of a class of models they’ve been working on for some time. They tout 8x lower costs and the opportunity to do on-premises deployment with this Mistral Medium series of models, and their tagline is “Medium is the new large.” So I guess, Chris, maybe I’ll start with you. We haven’t talked about Mistral on the show for quite a while, and I think one of the reasons I wanted to bring them back up again was obviously this new release, but it kind of offers the question: is Mistral still a contender in this open-source space? Curious about how you size that up.
Chris Hay: I think so. I love Mistral, first of all. And to come back to my earlier point, the Mistral team were the original folks that came up with the Llama models in the first place—Llama 1. So they are great innovators. I love the Mistral models. I especially love the Mistral 7B from a few years ago. And I think the new Mistral Medium is great, but maybe the criticism I would give is: have we had that sort of 7B or 8B model, or even a 3B type model from them? They’ve been focusing on... the Mistral Medium 3 is probably a 70-billion-parameter model. Even their small model they recently released was a 24-billion-parameter model. So when we think about the world that most open-source developers live in, it is the world of Llama, the world of Hugging Face, and therefore you need the smaller models. So by not having those smaller models out there, maybe they just sort of fell out of our thinking a little bit. But what they’re doing in the lab, Mistral Medium 3 is a fabulous model. It’s super fast; it really is state-of-the-art. But again, another thing I would criticize is they haven’t put a reasoning model with that as well. So it is as good as it is, but we’ve also moved on to reasoning models, and I think they just need to kind of push that. But I am confident they’re gonna do some great stuff, and we’re gonna see a resurgence of them.
Tim Hwang: Yeah, for sure. Kaoutar, I’m curious if you agree. I think Chris is making an interesting point, which is: “Medium” is maybe still too large for a lot of what’s happening in enterprise. It kind of sounds like, at least the story Chris is telling, is that they are kind of missing the boat on where the current competition is and where the current heat is in the open-source space. Do you think that becomes a problem for them going forward?
Kaoutar El Maghraoui: I think so, actually. Mistral’s strategy has been these open weights; they also focus on high performance and low inference costs. They have all been great, and while they haven’t been making many headlines like OpenAI or Anthropic right now, I think their consistency in releasing strong open models is worth noting. So the Medium 3, for example, ranks competitively on standard leaderboards. And I think their commitment to the open-source community is a very rare stance in today’s increasingly closed ecosystem. So the question is: has Mistral really faded from the race? Or are they quietly building the foundations for long-term impact in the open-weight AI ecosystem? Which I think they’re heading that way. Of course, like Chris mentioned, they need to fix the reasoning aspect of their models, but I think they’ll get there.
Tim Hwang: Yeah, for sure. Volkmar, one of the interesting parts about Mistral for me is that it feels like countries increasingly have their big AI champion, right? So there’s DeepSeek in China, and the US arguably has OpenAI and a number of companies. And for a long time, Mistral really was the big hope of Europe and certainly of France, on the idea that they would have a national champion that would be able to lift many boats and help build an AI industry and leadership for the continent. I’m curious if you buy that as a thesis. I know in your opening remarks you said you wish them all the best, which I guess the more critical way of saying that is you don’t know if they have a very strong hope. But curious if you wanna talk a little more about that.
Volkmar Uhlig: Okay, that’s a loaded question. If you look where GPUs get deployed, there’s a very strong concentration in the States and in a couple of countries which are trying—like we are talking about this with the Saudis—there’s investment going on. But if you look at the majority of GPU deployment, it still happens in very, very few places in the world. The 50,000-, 100,000-GPU clusters are pretty much only in the States. And so I think there is a concentration of capital and a concentration of skills, which is clearly not in Europe’s advantage. Europe is really good at writing regulations right now, but they’re regulating something which they cannot build themselves. And I think this is a real big danger. And so I think companies like... starting an AI company today in Europe is kind of insane, whereas it’s an under-understood market and already overregulated. And so I think Mistral is kind of walking the line. I think in general what we are going to see is that companies are going to focus on their local markets. If you look just from a language perspective, all these models are very English-focused, and then all the other languages are almost translations. And so the Chinese said, “Okay, we want to have a Chinese-first model.” And so I think there will be kind of local champions. I think also that we see a trend... you’re saying like “Medium is the new large,” I would say “small is the new medium.” And every six months that’s the case, right? So if you look at the capabilities of the models, what we could only do in huge models is now moving into something you can run on your laptop. But it took like two or three iterations. And so I think Mistral is just adopting to that general trend—that the technology is now at a point that I don’t need a 70B model anymore to get that type of performance. And it’s just the nature of the beast. I think also the smaller models are just much more economic, and so there is just economic pressure. The moment you deploy this at scale and you don’t have money to burn anymore, then you need to actually look at what the cost footprint is. And I think that’s probably a reaction of Mistral as well. You have to be close to where the clusters are to build the talent and expertise.
Tim Hwang: Yeah.
Volkmar Uhlig: I think the capital follows the talent; the talent follows the capital. And so you are in a world where Europe just doesn’t have deployments. So the people who are really good, they come here. I mean, you need to go somewhere where someone is willing to pay for 50,000 GPUs. That’s not happening in Europe.
Kaoutar El Maghraoui: Yeah. I second Volkmar’s point. Europe today, where it’s lagging is the compute infrastructure. The continent is really short on sovereign AI compute. There is no European equivalent of the AI 100/H100-scale clusters that you find in the US and even in China. There’s also a lack in VC funding; the startups in Europe today are struggling to access the scale of funding that fuels Silicon Valley and the Chinese tech ecosystems. These are things that are kind of slowing their progress in foundation model development. So it doesn’t have a direct equivalent to OpenAI or Google DeepMind. So Mistral, even if they’re champion... I think Volkmar pointed out, access to compute is very important. So we need to have both.
Tim Hwang: Chris, do you buy this? I mean, my note of skepticism is: we live in the 21st century, right? There’s the internet. The idea that you have to be proximate to all the compute in order to build a strong AI industry—I guess it works a little against my intuitions. But Chris, do you wanna jump in?
Chris Hay: I don’t buy any of that in the slightest. Have we all got short-term memories or something? What were we saying about China? Six months ago we were like, “Oh, US is the greatest.” Well, they don’t... China doesn’t have access to any H100s, and then DeepSeek comes along, and we all go, “Uh, well, okay. Whoops.” And I think that’s the exact same case. I mean, let’s look at what Mistral has done there. Right? We’re sitting arguing about “Medium,” but it’s one of the strongest models that is out there that’s non-reasoning. So they’ve shown that in the benchmarks. It is a great model if you go and use it. It is a fantastic model. And again, I’m guessing it’s a 70-billion-parameter model, but that’s phenomenal. What they’ve done is outperforming the Llama 3.2 Maverick model, for example. And we’re saying Europe is useless? Sorry. You just created one of the best models. And again, let’s go back to being Europe for a second. Who is leading Google’s Gemini model? Demis Hassabis. Where does he come from? Not America. If we look at OpenAI, who started that from an architectural point of view? Ilya Sutskever. Where did he come from? Not America. So there is lots of European innovation coming through. European companies are building stuff, and they don’t have access to AI clusters? Just... wait the first part, maybe not the second part. But most part...
Volkmar Uhlig: You just called out individuals. And if you look at the top companies founded in the Bay Area, that is most, like 50% or so, are not from Americans—like American-born. That’s a non-argument because what you are saying is like my first citizenship matters. It’s like, no, it matters where you actually build your company.
Chris Hay: Okay, but let’s take the Llama models, which were built by France, right? That was that team that was doing that. So innovation is coming from Europe.
Volkmar Uhlig: Hang on. Two things: businesses, is that innovation? Europe has the brainpower. Europe has no money to actually fund it, and that’s why it’s all funded in the United States. If you look at the amount of capital which is spent in the US on new technology and the amount of capital spent in Europe, there’s nothing spent in Europe. And therefore we—I’m very happy we got a full brain drain. Let’s take all these people and make it happen in the United States.
Chris Hay: But we literally had this argument six months ago about DeepSeek, right?
Volkmar Uhlig: So I’m telling you, that’s a different story. Completely different story.
Chris Hay: I don’t think it is.
Volkmar Uhlig: In one case you have a wall and people cannot get out, and the other case you don’t.
Kaoutar El Maghraoui: No, DeepSeek, they have the GPUs, but they probably didn’t have the latest, you know, or top-scale GPUs, but they were very clever in how... you know, take those limits and even act at the PTX level of NVIDIA, of the CUDA, so they can basically overcome the limitations that they have. So I think there are two things here we’re talking about: there is the brainpower, and there is the infrastructure and the deployments—the resource. The R&D, yeah, it could come anywhere. But then when you want to deploy these things at scale and see the business value, that’s I think where we see the lack.
Tim Hwang: So this is a very nice segue into our next segment, actually. It’s another story I wanted to bring up and have the panel react to, but I think it’ll be a continuation of this discussion in some ways. Super interesting news coming out of Saudi Arabia this week: NVIDIA announced that it will be collaborating with the Saudi investment funds to build what they call “AI factories.” And they are projecting the deployment of several hundred thousand advanced processors in Saudi Arabia over the next amount of time, and they’re promising a data center capacity of as much as, quote, “500 megawatts.” The first step of this is gonna be a shipment of 18,000 GB200 Grace Blackwell chips. Volkmar, maybe turn it to you first: can you give me a sense of what is 500 megawatts compared to where we are today?
Volkmar Uhlig: There are two viewpoints on this: that’s a lot, and it’s nothing. If you take a typical rack in a data center, it’s about 20 kilowatts. If you’re really trying to push the envelope, it’s 30, 35, but that’s kind of on the edge. So 20 kilowatts. If you take 500 megawatts, you do the math: it’s thousands of racks. Now, if you look on the flip side of what NVIDIA announced with the next-generation rack, which is a petaflop in a rack, it’s 600 kilowatts per rack. So if you take your 500 megawatts, then you can fit 800 of these racks in a data center. Now, 800 racks is still a very large installation, but it’s not 20,000 racks, right? So we are in a world where it’s a sufficiently large deployment to actually make a dent. For a whole country, that’s probably not enough.
Tim Hwang: Kaoutar, one of the things we were talking about earlier, I guess, was Volkmar’s thesis that the talent follows the capital. And so I guess maybe you could draw a comparison where you say, well, Europe’s got the talent, but it doesn’t have the capital. This seems to be a case where the Gulf states have the capital, and now the question is whether they’re gonna be able to bring the talent to really build out a much broader AI industry. I’m curious about how you think the prospects are of these kinds of deployments being the thing that allows you to trigger global leadership in the technology.
Kaoutar El Maghraoui: Yeah, that’s a very interesting point here. It’s like what we’re seeing in Europe: talent is there; the capital is lagging. I think what’s happening in Saudi Arabia—the deal that’s happening with NVIDIA and AMD and the US—it’s marking a major milestone in the rise of what we call sovereign AI infrastructure. So Saudi Arabia here is not just buying chips; they’re taking a stake and trying to claim their place in the future of compute. And with the Grace Blackwell deployments, they’re trying to also commit to Arabic language LLMs and nearly, I think, about two exaflops of AI power on the horizon. So it’s not just a regional experiment; I think it’s a global power play that they’re trying to play here. And I think from the US side, this is also reflecting a clear shift in the AI export strategy—where China is restricted, Saudi Arabia and the Gulf nations are the US’s preferred strategic partners for high-end AI. But the issue here is: these sovereign AI efforts need not just the silicon, not more than just the compute. They’ll need also the open ecosystem, the talent, the responsible development frameworks to truly compete. So I think we’ll have to see what Saudi Arabia is doing. They’re getting this huge infrastructure, building this huge infrastructure, but the ecosystem and the talent, they still need to work on that.
Tim Hwang: Well, Chris, maybe I’ll bring you in. I think you were a Europe booster, and I’m curious about if you think you are a Gulf States booster now, given seeing these investments. Will they be able to bring the talent in the future? Will a European researcher say, “Well, I could go work for a company in the US, or I could go work for a company in the Gulf”? Like, do you think those dynamics will start to happen as we see these deployments get bigger and bigger in, say, Saudi Arabia?
Chris Hay: I think they already have the talent. So I’ve spent quite a bit of time in Saudi Arabia with a bunch of these companies, and they’re heavily investing in AI. And again, even if we go back a couple of years, just after the Llama models came out—the first ones—what was the most popular model at that point? It was the Falcon models. And where did they come from? Saudi Arabia. So they already have the talent within the region, between Saudi and UAE, who have been developing some of these models in-region. And I think therefore we’re gonna see talent flock towards that infrastructure, to Volkmar’s point earlier. So I think you’re gonna start to see stuff coming out of them. And back to my earlier point: even with that amount of infrastructure, when you place constraints, people get more creative. So I think they’re gonna do interesting things. And then again, maybe taking language-first models as well, you’re gonna start to get different flavors as opposed to an English-first model—again, it was brought up earlier. So I think we’re gonna start to see different flavors of models that may have different and newer capabilities, which will become a plus to the overall world ecosystem of models. So I am positive on that, and I think it’s a good thing.
Volkmar Uhlig: If you look 20 years back, we had a similar investment cycle, and this was affected by the massive build-out of internet infrastructure: fiber optics in the ground, data centers, computers connected. And so I think there’s a certain repeat—like those nations, because they have a different investment philosophy, like look at Singapore. Singapore made the decision: “We want to be the data center hub and the fiber optic hub of that region,” and they poured billions of dollars into it, and that attracted a lot of business. I think those nations are, because of a central command structure, very advantaged in pushing those types of large-scale infrastructure investments out of a sovereign wealth fund and saying, “Okay, we put this capital to work.” They usually are more challenged if it’s a pure IP play—like, okay, we need to build some software—because then they don’t have a competitive advantage. So I think they’re all playing to their strength here, saying, “Okay, we take oil money and convert it into GPUs,” and thereby creating that suction sound of getting people into the region. So I think it’s a very natural play for those types of regional players of trying to establish dominance in a new emerging field, which has this—oh, and by the way, we need to put 20, 30, USD 40 billion down, which you don’t do out of normal private funding; you need someone like the US government scale. It’s government scale, and that’s exactly where they can shine, and that’s where you’re seeing it.
Tim Hwang: Exactly.
Volkmar Uhlig: It’s government scale, and that’s exactly where they can shine, and that’s where you’re seeing it.
Tim Hwang: Yeah, for sure. And maybe a final question on this before we move to the next topic: Volkmar, as someone who’s very much in the infrastructure game, my friend was making the argument to me recently that, in some ways, actually it’s possible that Saudi Arabia might be advantaged even against the US on data center build-out because of, for example, the ability to access and move energy assets around—something that they’re not gonna face the same kind of permitting issues and construction issues that you have in the US. Do you buy that as an advantage that they have?
Volkmar Uhlig: No, I don’t.
Tim Hwang: Okay.
Volkmar Uhlig: Because we have 50 states, and I moved out of California to Texas, and it’s totally different; permitting is so much easier. So I think we will see a similar thing: they’re playing it on a country level; we will play this on a regional level. And if you look at the US, there is pretty much just a network ring which goes all through pretty much all states. And so you can put your data centers in Arizona, Oregon, Texas—you go where the power is cheap for these types of build-outs.
Tim Hwang: Yeah. I like that take. It’s like we have Saudi—it’s called Texas.
Volkmar Uhlig: Exactly. We also have the oil.
Tim Hwang: Yeah, right. Exactly. Well, great. I’m gonna move us on to our next topic, moving us a little away from the world of chips and national competition to something a little more close to home. Two interesting releases that happened fairly recently: one of them was from OpenAI, this benchmark they released called Health Bench, which is specifically a curated set of about 5,000 conversations of interactions between AI models and users or clinicians, and the idea is to create standard benchmarks for AI’s use in the health domain. There’s also a really interesting benchmark that came out of IBM called IT Bench, which is looking at benchmarking on agents. Kaoutar, maybe I’ll turn it to you. The funny thing that I have when a new model releases now, whether it be Mistral Medium or what have you, is that they always say, “We’re very good against all these benchmarks,” and then there’s a list of like 50 benchmarks, and it’s very difficult to tell what it actually means. There’s lots and lots of benchmarks, but it’s very difficult to say, “Okay, I’m gonna deploy this in the health space; is this actually a good model for me?” And I’m curious: I wanted to put to you the idea that in the future, these benchmarks might end up becoming a lot more fragmented than they are right now, where you might imagine we say, “Oh, we’re gonna release a model, but it’s gonna be specifically for health applications, and here are the benchmarks for it.” Do you think that’s where we’re headed, or are we gonna keep developing out this ever more comprehensive benchmark suite for every single model that comes out?
Kaoutar El Maghraoui: Yeah. I think this race towards building these AI models and benchmarking against them is gonna continue. But now that we’re entering the age of agent AI, the benchmarking still seems like it’s in the chatbot era. So if you want to deploy trustworthy AI agents, I think we need new evaluation frameworks that combine general reasoning metrics with domain task completion. Think of it like: general benchmark tests the IQ, but sector-specific ones test the job performance. So I think ultimately the future of benchmarking lies in these hybrid evaluation stacks: general foundation models tested across standard reasoning tasks, but also we need stress tests in realistic operational settings—like the examples you mentioned from OpenAI, the Health Bench, or the IBM IT Bench. Those are very domain-specific, and we need more of those. So we need the general evaluation stacks and frameworks, but we also need the stress tests for these realistic operational settings. And those will become, I think, industry standards, not because they’re broad, but because they’re really real. So the new wave we’re seeing with OpenAI’s Health Bench and IBM’s IT agent benchmark—these really are raising critical questions: are we really measuring the right things? And as I see these models shifting from static chatbots to dynamic agents, the traditional benchmarks need to change.
Tim Hwang: Yeah. What I like about that—and Chris, I’m curious about your comment on this—is: does the idea of “state-of-the-art” even make sense anymore? It kind of feels like there are so many use cases now for AI that it’ll be very difficult to imagine that one model is state-of-the-art across all applications. And so are we kind of maturing our thinking here? Does it... I don’t know, Chris, do you buy that “state-of-the-art” actually doesn’t make any sense because of the number of applications?
Chris Hay: I don’t think “state-of-the-art” makes any sense. I think benchmarks—it’s all marketing hype, isn’t it? “We are the greatest at this, and look, we are topping this chart here.”
Tim Hwang: Yeah, you need the chart that shows that you’re ahead on all the benchmarks. That’s what you do.
Chris Hay: Exactly. Every single model provider releases their chart, and they beat everybody else. And they select the models that they don’t want to pit against. So it’s like, “I’m selecting this, this, and this. Look, I lead on all of this, and I win at this benchmark on that one,” et cetera, because it’s gotta look like you’ve got the greatest model ever. And I understand that, but I truly think: don’t let somebody else tell you what model is good for your use case. So if you need a benchmark, go create a benchmark for yourself. I’m in this domain; these are the sort of things I’m gonna do; I’m gonna test for it; and I’ll create my own evals, and then I’ll make sure that I’m using that model for the purpose and tasks that I want. Now, don’t get me wrong, benchmarks are kind of useful in some regards because that allows you to basically go, “Yeah, this model’s probably around the right level,” and therefore I can go and try it on this different stuff. So it gives you an idea of whether it’s something you’re gonna be useful for. But the reality is most of us know what model is good just from using it. So I’m a vibe-er all the way.
Tim Hwang: Well, but Chris, when you say you create your own benchmark, isn’t that a biased approach? I mean, I can craft it in a way that’s gonna show my model as the best in whatever I’m doing, so there might be some bias there, as opposed to having an external party create maybe a set of benchmarks.
Chris Hay: When I say create my own benchmark, I mean for my specific task. So if I’m creating a chatbot for tourism that answers FAQ questions on flights, right? You may have the best coding model in the world, but if it’s giving me flights for another company, then it’s not really good. Or if it’s not reading the FAQ, then it’s not really any good. So like any normal software application, I’m gonna want to create test cases to say, “Is this thing coming back with the answers that I’m expecting for my business use case?” That’s the kind of eval. So it’s not about testing for generality; it’s actually “Is the model any good at the task that I want it to perform?” And rather than relying on somebody else to tell you whether it’s any good for your application you’re building, go create some evals and test it out for yourself. And most of the time, back to the point, a smaller model that is fine-tuned or designed for a specific application will do better than a general model anyway.
Tim Hwang: Volkmar, I’m wondering if you have any predictions on where the meta of this moves over time? ‘Cause I guess the world Chris is describing, which I really agree with, is everybody sees these charts of state-of-the-art performance, and yeah, I think the main thing that I get from them now is, “Okay, you did your homework; you’re at least as good as everyone else.” It doesn’t necessarily make me stand up and be like, “Wow, this is the most incredible model ever,” but it just says like you’re doing as good as everyone else. If that’s the case, then we’re living in a world where these kind of marketing statements don’t really have that much value anymore. And I’m kind of curious how in the future an OpenAI or a Mistral or whoever is gonna demonstrate that, “Oh, this is the model you really should be using,” if not these benchmarks.
Volkmar Uhlig: So I’m very much aligned with Chris what he said. I think there is the outside marketing: I wanna bring my product to market; I need to communicate what it does. And I think what we will see is—and this is what you get with the medical benchmark—first it was, “How good are you on math tests and physics tests and logic tests?” And now what’s happening as we are widening the use cases, every industry will come and say, “Hey, I’m the medical guy, and I’m the rocketship guy. I wanna know how good that model works.” In the end, when you are putting things in production, you build your own regression tests, because what happens is that you have something that fails, it goes into regression test. And then what I’m seeing with projects we’re running internally here: every six months you upgrade your model. So if you don’t build your regression test, you actually don’t know what’s going to fail when you’re switching from the old model to a new model, and then you have a problem with the customers. And so over time, you are building effectively a benchmark—call it benchmark, call it regression test, call it unit test—it doesn’t matter, but you’re building out your way of validating that your fine-tuned model or the next version of the open-source model you are using actually works for your use case. And if it doesn’t work, you cannot upgrade. It’s a quality assurance thing. And that quality assurance is extremely domain-specific.
Tim Hwang: That’s right. Yeah. It almost suggests a world where there’ll be a lot more eval work that needs to happen in-house. And then maybe finally, people have been talking about this for years, but maybe actually finally creates a market for specialized businesses that just do evals, because you live in a world where basically your big foundation model company’s not gonna run every single eval in the whole world, but if you have a specific use case, you need someone with eval expertise. It almost feels like that suddenly starts to become viable in this world as the market matures.
Volkmar Uhlig: But I wanna bring it back to Kaoutar’s point, and I think it’s really relevant, which is we focus on the model world. But as Kaoutar was saying, in a world of agents, it’s gonna be less about whether this model is capable of performing this task; it’s gonna be more about is the agent capable of doing this task and how good is it performing?
Chris Hay: So...
Volkmar Uhlig: So I sort of agree with that point, that the benchmarks kind of need to move on and really look at it from a different perspective as opposed to just always looking at the models, which I think will also give a better quality because agents usually have a much more narrow band. This is the problem with the generic model: I have OpenAI; it can do anything. And so how do you test that it does what you want to do specifically? But once you go to agent, you really narrow the applicability of the model because now you’re saying, “Agent, you do flight booking and nothing else.” And now I can go deep, not just broad. Right now we just go so broad: a little bit of math, a little bit of medical. But now if you have an agent, you can actually say, “Do your task, or you don’t.”
Tim Hwang: I’m gonna move us on to our final story in the last few minutes of the episode. This is kind of a fun one. On MoE we cover infrastructure and chips and health and model evals; we don’t really talk about showbiz all that much. But there was this interesting story: Amazon was doing its upfront event where it talks to advertisers for Prime Video’s upcoming season, and there was a bunch of announcements—new shows they’re doing, all the usual show business stuff. But there was one interesting AI hook that was mentioned in this news story where Amazon announced that they were gonna start using generative AI to create contextual advertising on Prime Video. They didn’t really provide too many details, but the idea would be that you’d be watching a show on Prime Video, an ad would come on, but it would be an ad generated on the fly using AI, and in a way that’s presumably contextually related to both what you’re watching and what it knows about you. Which is a really weird, interesting change. We’ve had targeted advertising for many years, of course, but this seems to be a qualitative shift where the actual ad itself will be generated on the fly. Chris, are you a fan of this? Are you excited about this?
Chris Hay: I am not excited about this whatsoever. I mean, come on. It’s a loaded question, but yeah.
Tim Hwang: Chris, go ahead.
Chris Hay: No, seriously. I mean, how many times on Amazon—like I look at, I don’t know, maybe I buy a comb off of Amazon, and then for the next three weeks I am seeing combs appear everywhere. Every website I click on, I buy a comb. Everything I look on Amazon: “Do you want a comb? Here’s a comb, here’s a comb.” I’m like, “I just bought a comb. Don’t stop trying to sell me combs.” The last thing I want to do is stick on the TV, there I am about to watch some New York Giants be terrible, and then guess what? More combs arrive. You’re like, “Oh, great. Oh no.” So no, I... once you can stop the combs following me around the internet, then I will be excited about Gen AI adverts.
Tim Hwang: Okay. Quick hit, Kaoutar: are you excited by this?
Kaoutar El Maghraoui: Yeah. I’m also worried about this. I think this contextualized advertising might be a little too much—this hyper-personalized advertising. I’m not sure; I think we’ll have to see how the viewers react to this. I think we’ve seen clearly Chris’s reaction. And also it’s machine-generated content inside the show. So yeah, I’m also a bit worried about this experience and how it’s gonna play. So I feel it’s just like they’re more getting into our minds and what we see, and so it’s a bit creepy.
Tim Hwang: And last but not least, to close out the episode, Volkmar, your take on this incredible new feature that Amazon’s about to launch into our lives.
Volkmar Uhlig: So, two startups before I had an ed-tech company, so I probably know too much about it. There’s a big challenge of creating really good creatives, and I think that is already now at a point where AI is really helping—not just internal clicking together or something, but actually creating really good advertising. It’s going to be really interesting to see: they can now do native formats. So native formats in print is like when you have something that looks like it’s part of the article, and so you are reading it despite that it’s an advertising. And so what you now could do is you could go full native: you have the movie, and you could even plug something; you could change the plot of the movie if you take it to an extreme. So it’s going to be interesting to see. I wanna see all the issues they have there: the model accidentally renders something you shouldn’t render, so how do you do quality control? But I think this is the path where we are going. Image rendering is obvious, but video rendering is the next step. I think it’s hopefully a long path until this becomes reality. Hopefully.
Tim Hwang: Yeah, we’ll see. I just envision that you’re like watching Star Wars in 2030 and Luke Skywalker’s like, “You should really buy a comb.”
Volkmar Uhlig: Exactly.
Tim Hwang: That’s not a good outcome for sure.
Chris Hay: Comb scene.
Tim Hwang: Yeah, everybody remembers that one. So, well, that’s all the time that we have for today. Kaoutar, Volkmar, Chris—pleasure always to have you on the show. This is one of my favorite panels on MoE. And thanks to all you listeners for joining us. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify, and podcast platforms everywhere. And we’ll see you next week on Mixture of Experts.
Discover how to harness AI to drive business growth and innovation with insights from industry leaders and practical strategies.
Applications and devices equipped with AI can see and identify objects. They can understand and respond to human language. They can learn from new information and experience. But what is AI?
It has become a fundamental deep learning technique, particularly in the training process of foundation models used for generative AI. But what is fine-tuning and how does it work?
Listen to engaging discussions with tech leaders. Watch the latest episodes.