OpenAI social network, Anthropic’s reasoning study and humanoid half-marathon

Watch the episode
Mixture of Experts album cover
Episode 52: OpenAI social network, Anthropic’s reasoning study and humanoid half-marathon

Is OpenAI going to enter the social media game? In episode 52 of Mixture of Experts, Gabe Goodhart, Kate Soule and Marina Danilevsky join host Tim Hwang. First, Sam Altman is rumored to be testing an internal prototype social network; why is this a potential next move for the AI giant? Next, for our paper of the week, we analyze Anthropic’s study on chain-of-thought reasoning, “Reasoning models don’t always say what they think.” Then, AI scraping puts a strain on Wikimedia; what’s the impact of this? Finally, China held a humanoid robot half-marathon, where humans raced alongside robot competitors. Who wins this AI race? All that and more on today’s Mixture of Experts.


Key takeaways:
  • 00:41 – OpenAI social network
  • 10:02 – Anthropic's reasoning study
  • 20:56 – AI bots strain Wikimedia
  • 31:33 – Humanoid half-marathon
Listen on Apple podcasts Spotify podcasts YouTube Casted

Episode transcript

Tim Hwang: OpenAI is apparently working on a new social network. Pretty cool or kind of cringe? Kate Soule is Director of Technical Product Management for Granite. Kate, welcome back to the show. What do you think?

Kate Soule: Major cringe vibes. No, thank you.

Tim Hwang: Okay. Marina Danilevsky is a Senior Research Scientist. Marina, cool or cringe?

Marina Danilevsky: Extremely cringe.

Tim Hwang: Okay, I’m gonna have a unanimous vote on this one. And last but not least, Gabe Goodhart is joining us for the very first time, Chief Architect, AI Open Innovation. Gabe, welcome to the show. What do you think?

Gabe Goodhart: So many things that could go wrong. Maybe something interesting, but cringe for me as well.

Tim Hwang: Okay, great. We’ll get into that, all that and more on today’s Mixture of Experts. I am Tim Hwang, and welcome to Mixture of Experts. Each week, MoE brings together the sharpest crew in all of podcasting to discuss and debate the biggest news in artificial intelligence. As always, there’s a lot to cover. We’re gonna talk about a super interesting blog post out of Anthropic about reasoning models, Wikipedia getting slammed by scraping bots, and a super interesting half marathon being run by robots. But first, I want to start with the round-the-horn question that we began with, which is rumors that OpenAI is going to launch its own social network, which of course is baffling as a company that’s largely built its money, its expertise, and its brand on foundation models and advancing the state of the art. Maybe, Kate, I’ll turn to you first. Why would OpenAI wanna do this at all?

Kate Soule: Yeah, I actually don’t think it’s that baffling. I think it’s pretty straightforward. I mean, Meta and X both have these social platforms that they can use to learn about conversational patterns, to frankly generate and collect data potentially. And OpenAI and other providers have shared that they’re running out of data, so to speak. And so I think they very much see this as a data play of being able to create a platform. Hopefully they have some way to provide value to incentivize users to join, but ultimately I think they’re in it for the data that they’ll be able to collect behind the scenes and use that to train more conversational, fluent, and robust models in the future.

Tim Hwang: Got it. Yeah, and I think, Marina, that was a question I had for you: is this social media data actually all that valuable? I go on social media and scroll on X, former Twitter, and I’m like, this is kind of like garbage content in a lot of ways. But is this data actually helpful in advancing the state of the models and what they can do? I think there’s an interesting question about this. OpenAI clearly sees some upside, but I’m curious about what you think.

Marina Danilevsky: I mean, a little bit of it I think is FOMO of, “Wait, we want people to come and make bad internet memes on our platform. Why do they have to leave our platform? We wanna be there too.” But yeah, any of this kind of data is valuable because it’s different. So again, synthetic data generation, which is where everybody is really getting their data now that they’ve run out of data, isn’t really good at making interesting little viral memes. These models aren’t that great with humor and subtlety and creativity and things like that, so you get that from people. So, especially being able to combine this type of additional input and injection of ideas that you would get from this kind of thing... Yeah, I will say you’re probably gonna get a real specific slice of humanity using this... a real specific... Yeah, right. No comment... that you’re gonna actually have using this and creating the data. So yeah, you’ll get something out of it. And I agree with Kate, as I usually do, that it’s a data play. And also, yeah, it’s not that hard anymore to put this together. But again, I think they’re gonna be a little limited in who comes there and uses it for what.

Tim Hwang: Yeah. And Gabe, I know you were maybe the one... everybody thought it was cringe, so maybe that’s just an established fact, but you were saying like it might be cool if they maybe get a couple things right. What do you have in mind there?

Gabe Goodhart: Yeah. Well, I think the part that’s really interesting for me is thinking about this as a way to experiment with a novel interaction pattern. Personally, social networking seems to me to be the wrong way. They’ve already pioneered a fairly novel interaction pattern, but I think where many AI platforms are finding themselves hitting a wall is integrating further into daily life as opposed to bringing the users to the AI. So I think to me, this seems like sort of a, “Well, where do people interact? Hmm, I wonder if we could go there,” type of play. And obviously as a company that’s trying to make themselves a relevant standalone company as opposed to an integration partner, they’re not gonna look to plug their AI directly into a competitor’s platform, since many of the other social media companies are also AI model companies at this point. And so I’m wondering if this is really their attempt to branch out from being an AI-specific company to being a tech company, right? Think Google moving beyond search, think Meta moving beyond Facebook. I’m wondering if this is just their first step in trying to become an omni-company like the other big tech players that are user-facing.

Tim Hwang: Yeah, that’s an interesting flip. And I think it introduces a new argument, right? On one hand, there’s sort of Kate and Marina’s take, which is, “We just need the data for the models.”

Gabe Goodhart: Oh, and data is a hundred percent part of the story, no question.

Tim Hwang: Totally. Yeah.

Gabe Goodhart: Yeah. And I’m not denying that. I mean, I think that makes total sense.

Tim Hwang: And then the second bit that you’re bringing in, which is sort of interesting, is like, well, it’s also maybe if it gets popular, a distribution point for AI models, which is a funny thing. The phrase “social media” implies people talking to people. And so it’s interesting, the idea that, okay, well actually we really want to augment that or get AIs involved. It kind of reminds me of the craze about, oh, you’re gonna have a group chat, but then there’s gonna be an AI in the group chat that will assist. And by and large, those haven’t taken off all that much. But maybe the nature of what we think of as a social network is changing. I don’t know. Kate, if you think that’s a possibility that we see happening here.

Kate Soule: I don’t see a ton of future there, but what I do think they’re setting themselves up for is another form of monetization and integrating with advertisements—so having models that innately understand the language being used to communicate with one another on their platform and using those to then generate really targeted advertisements directly to those users is something that OpenAI is really setting themselves up for. And so, as they’re also looking to pay for all of those really expensive models, they probably see some opportunity to have new sources of revenue based off of this new platform.

Marina Danilevsky: Yeah, to jump on that... when Facebook got started, it was a really big deal for advertising because you could actually do things that were super targeted, really for the first time in that sense, on that kind of scale. So rather than general, “Oh, this is what you searched for,” now it’s, “No, we know who you are, what you like, what your friends like, all these characteristics about you.” This is trying to see if you can take it to the next level, as well as what your speaking patterns are. I mean, imagine... I don’t think it’s gonna be humans socializing with humans. I think it’s going to be bots selling to humans using their exact speaking patterns, languages, everything that you learn in sales about trying to, like, body language mimic the person you’re talking to is gonna be tenfold with these models. Like personal influencers, right? You’re now gonna get an influencer that’s directed straight at you. So there’s just no buffer left between you and capitalism, just direct pipeline.

Tim Hwang: Marina, do you have a point of view on this?

Marina Danilevsky: Just a little.

Tim Hwang: Maybe a final thing to talk about before we move on to the next topic is, to take a step back, we can almost think about this outside of OpenAI kind of working on this fun, weird thing. It feels like if you buy this sort of data interpretation, we’re gonna see all sorts of weird acquisitions happening going forward, right? AI companies in their hunger for data will acquire and launch all sorts of services largely for the data value. I’m still looking for, is an AI company gonna buy a law firm at some point because it has a bunch of data about how lawyers interact? That’s really valuable for training these models. Another one that people have talked about is acquire a call center, right? ‘Cause you really want all that customer support interaction. Curious if the panel has views on other places where this could go, as almost like the drive to train models motivates all of this vertical integration that maybe we find a little dissonant when you hear about OpenAI doing a social network.

Gabe Goodhart: The other take I had when reading about this story was that it was an attempt to break out of model commoditization. So I think other models have caught up to OpenAI frankly, and marginal gains in quality are really not driving the use case anymore. So on the one hand, having a captive platform that gets users into their platform is really valuable from a moat perspective. On the other hand, the data also helps them have a source of data that is hypothetically differentiated and lets their models actually reach a capability that other models can’t. So I think you’re spot on with this idea of differentiation based on upstream data sources that are owned and completely enclosed by the model authors.

Marina Danilevsky: Upstream data sources and downstream use cases as well, right? So the more we go into multimodality and models can do this, that, or the next thing, then there’s a desire to say, “Okay, great. Instead of focusing now, we’re gonna go ahead and diversify.” Again, nothing new under the sun. This puts me in mind of historically when you had oil companies earlier in the thirties, forties, fifties that were like, “We’re gonna buy movie studios.” Why? Because economically, we want a stock portfolio that’s diversified. It’s weird to us now, but it’s a little bit of that same repetition of, “Hey, can our stuff actually make it everywhere up and down?” So maybe this is just the newest version of that particular wave.

Tim Hwang: All right, well, we’ll hang on tight. Everybody agrees that it is cringe, but they will take their best possible shot at it. I’m sure we’ll talk about it when it finally launches, if it does launch. So I’m gonna move us on to our next topic. Super fun blog post coming out of Anthropic, building on some of the research they’ve talked about in the past. What I liked about it was it sort of brought up an issue very crisply that I think is worth talking a little about. Basically, the blog post takes a look at reasoning models and whether or not the reasoning that models give for how they rendered a decision is sort of faithful. The way the researchers investigate this is the idea that we’re gonna give the model a little hint on how to solve a problem, and then we basically say, does the model disclose that it had this kind of unfair hint when it accomplishes a task? They have this fun result: they say that Claude 3.7 mentions this hint only about 25% of the time. Their DeepSeek comparison is that it mentions it around 39% of the time. So the claim they’re making is that reasoning models often don’t fully expose all of the things that they do to make decisions. And, Marina, maybe I’ll kick it to you first. Talking to friends of mine, there’s a lot of hope that reasoning is this great interpretive tool that allows us to work with these models better over time. But this seems to cast some doubt. And you’re already shaking your head, so maybe I’ll just let you rant for a bit.

Marina Danilevsky: No, yes, I can rant for a while on this. This is my soapbox. I’m sure I’ve made this point on this show before: this reasoning isn’t real reasoning in the sense that we think of as reasoning, mathematical reasoning. I have been studying this paper all over the place, by the way. I love it. I think it’s really great. As you said, how they really crisply were trying to show... because it’s pretty hard to insert yourself into the model and say, “Well, what actually happened?” We started noticing almost immediately when the reasoning models come out, and you’re like, “Yes, but what happens when you have the answer?” And it’s mentioning things that weren’t in the reasoning. It’s an immediate red flag that there’s something else going on here, and this is a very nice way of being able to actually at least poke and have a little bit of local approximation of, “Well, are you even paying attention to this piece of information or not?” We see similar stuff when we even just try to do evaluations of faithfulness in general that’s content-based. If it is something very niche, the model may be like, “Look, I can’t even figure this out. I’m gonna fall back on things where I’ve got higher probabilities,” and go in that direction. So again, I like this kind of work. I like this kind of traceability of... will you think that it’s reasoning just because we gave it that name? It’s not. It’s yet another problem, like with words like “hallucination,” where you put an anthropomorphized word there and it means all these things that it does not mean. So I like this work. More of this, please.

Tim Hwang: Yeah, for sure. I guess what I’m left with is, Kate, maybe you have a take on this: so what is reasoning anyway? It appears to give a step-by-step disclosure or audit of how a model reached a decision. But for Marina’s rant, it’s unclear if it actually gives you anything. Is it just theater? What is it exactly?

Kate Soule: I think falling on Marina’s very well-articulated point, “reasoning” is a very anthropomorphized term. And when we talk about reasoning in the model context, what we’re really talking about is the model has been trained to generate more tokens before it makes a final answer. And that process has all sorts of reinforcement learning added on top to try and basically bring the model into a part of its distribution that will be more successful in making the final answer. And what I think the paper and the blog do really well is help articulate that chain-of-thought reasoning is not a proxy for explainability. So just because the model is saying X, Y, and Z in its chain-of-thought, it does not actually mean that the model thought through step-by-step. Where I do have a bit of a bone to pick with the paper and the blog in general is they themselves then fall into the trap throughout it of talking about the model is disingenuous, the model is deceitful, the model’s doing all these things. I think it’s very important to look at the paper and see the experiment that they ran. They injected an answer into the conversation history as if the model itself came up with that answer, more or less, from what I could tell. And then they asked the question and they saw, did the model refer to that previous answer? The model has not been asked to cite its sources, so to speak, in that context. The model has not been extensively trained... this is 3.7, I think is the first reasoning model that Anthropic put out, so it’s pretty early in their journey for reasoning. The model has not explicitly been trained to prioritize, “If somebody tells you an answer, make sure you cite that answer down the road.” There’s a lot of broader things you would wanna look at to make statements about being deceitful or disingenuous or even hallucinating. In a lot of ways, I think we’re just testing. This is one very narrow experiment, and it helps bring to light: don’t treat chain-of-thought reasoning as an explanation. I don’t think it’s fair to say that all chain-of-thought reasoning is false, or that the model only cited the hint 25% of the time, therefore it’s not leveraging chain-of-thought reasoning correctly to drive a final decision. I think there’s still a lot more work ahead.

Tim Hwang: Yeah, for sure. This thinking is so interesting, and I think it’s a great example of where going anthropomorphic on this is bad. Because in normal life, I’m talking to Gabe and I give some reasoning, and it actually gives you some explainability. But here’s a weird result where the appearance of reasoning helps the model get to the right answer, but it’s actually not an explanation, which is very weird.

Gabe Goodhart: As a result, these are all research hacks trying to boost metrics, right? And the metrics they’re trying to boost are often mathematical exams—that is where this discipline has evolved from. And so trying to ascribe it much more important meaning than what it was developed for is really dangerous. And it’s still very early on in reasoning for LLMs in general.

Tim Hwang: Yeah, for sure. Gabe, can you save us? Do you wanna propose, if not reasoning, what should we call this thing that the models are doing?

Gabe Goodhart: Yeah, I mean... I don’t know that I can save us, but you know, I wanted to say, “Hey Kate, what’s two plus two? By the way, the answer is two, please explain your reasoning,” and what’s Kate gonna say? Right? She’s gonna say “two,” and she’s probably not gonna say, “You told me the answer, that’s my reasoning.” So I think I really agree with everything you said, Kate, and I felt exactly the same quibble reading the paper about the way they anthropomorphize the problem. And the one that really stuck out to me is that the entire framing was that they were trying to discover the model’s internal reasoning process, and just exactly that phrase felt really wrong to me. And Tim, you said we’re having a conversation and I might explain my thought process, and that might help with explainability, but inside my brain, theoretically at least, there are lots of neurons firing that are not coming out my mouth. That’s not true of a model, right? A model does have its weights connecting to one another, doing matrix math, and we’re not looking at the specific weights of those, but the only tokens that are actually getting generated are the ones you’re seeing. And so to me, I think you said it exactly right, Kate. The chain-of-thought is a way of basically priming the pump in probability space so that the final answer is more accurate, and it’s completely mirroring the pattern of how it was trained. And so it’s useful from a human interface perspective, not useful from an actually unboxing what’s happening inside the math perspective. So it’s still a really cool trick, and it really helps because one of the real novelties of generative AI is that it’s speaking directly to humans. We think about pre-gen AI models, and their job was to encode something so that a programmer could consume it in a structured output form and then write a fancy program around it. Generative AI is speaking directly back in a human interface in a modality that a human can consume. And from that perspective, from a lay perspective, it’s really valuable to have additional words that help the human understand the answer, but it’s not necessarily ascribing an actual thought process, so to speak, to this pile of matrix math.

Marina Danilevsky: I would propose we can call it like “warm-up,” like you warm up before a sports event where you’re doing exercises, and that’s not exactly what you’re gonna do in the event, but it makes you better at the event itself. Or like warming your car up. And so you get these signals of what you’ve done during your warmup to, as Kate very well put it, prepare to give a better answer. But that’s what these are signals of. It’s not reasoning, it’s warming up.

Tim Hwang: I love it. Marina saved us. I guess maybe a final question and then we can move on to the next topic is... Gabe, I think you brought a really good perspective on the lay user of these tools. And one thing I’m left with with work like this is, should big model companies be exposing reasoning traces to users? ‘Cause it feels like the tendency is that people will read into it that it is literally how the model is making decisions, which is at least a little bit deceptive. It drives trust with the user, but maybe in a way that’s unwarranted. I don’t know what people think about that.

Gabe Goodhart: Yeah, no, it’s a great question, and I think you brought in the word “trust,” which is such an important and fuzzy word in this space. And again, Kate, you pointed out that so much of how we define these terms from a technical perspective is driven by specific benchmarks and specific problems we’re trying to solve. But at the end of the day, trust is about the consumer’s interpretation of their experience with the AI system. And I do think there’s some value in exposing a longer-form output in the same way that if you read an article written by a human and the article exposed the research collection process that went into creating it, you would have more trust in the output of the conclusion versus just presenting a conclusion. So I think there’s some potential value there, just purely from a human interpretability perspective. But I do think you’re exactly right, and Marina, I’m gonna lean on that warmup thing now. I love that framing. It really is all about warming up for the final answer that you’re gonna give, and not about anthropomorphizing some kind of thought process.

Tim Hwang: So moving on to our next topic, very interesting news story that was reported by Ars Technica. Basically, the Wikimedia Foundation, which runs Wikipedia as well as a number of other open knowledge projects online, cited a stat that was quite interesting, pretty shocking in some ways, that since January 2024, they have seen a 50% increase in bandwidth consumed on their service. And they attribute this basically to the rise of bots attempting to scrape media content from Wikipedia, largely trying to scrape data for the purpose of training AI models. This problem has gotten so bad that they actually more recently released a dataset on Kaggle in an effort to dissuade bots from scraping their site, to say like, “This is a nicely formatted dataset, you should use this instead.” We’ll see how effective that is. This is an interesting story in part because I think it goes back to the first topic we were talking about, which is these weird second-order effects we’re seeing as AI companies chase the dream of advancing their models. Kate, I guess I don’t know if you have any feelings about this. In some ways I’m a little protective of Wikipedia. I’m like, they should be blocking all the bots. We need to preserve the sustainability of these services. But it also means that Wikipedia is incredibly in high demand. And so I’m curious about how you navigate how we should feel about this.

Kate Soule: So I think a couple of things to make clear: Wikipedia is a very high-quality data source that has all of these really rich links describing how topics relate to one another. So that is an incredibly rich, valuable dataset for model training. But I think it kind of brings out an issue that’s more broadly being felt across all sorts of different content providers, which is on crawling, and particularly crawling that does not adhere to the guidelines and rules of the road that have been established, at least in the United States—things like chatbots or crawlers ignoring robots.txt files and other behaviors and practices that are starting to get a little predatory, essentially passing a lot of the costs that model trainers are going after onto the actual data providers themselves. So not only are they giving their data away for free because it’s available under fair use if it’s crawled in the United States, but now they’re also incurring additional costs. And we really need to, as an industry, have a broader discussion on how to responsibly engage with providers like Wikipedia. And it sounds like they’re setting up a number of those types of discussions, which is really exciting to see, to help not only have some of the more strict rules on things like robots.txt, but also have community-agreed-upon and defined best practices that we can use to more broadly enhance Wikipedia’s mission of sharing this data publicly without penalizing the content providers.

Tim Hwang: Yeah, this is one of the really interesting questions: how quickly can we get that balanced? One of my worries is a little bit of what happened with Reddit. It was a for-profit company, but the way I understand it is, well, actually AI companies really wanted to scrape that data. And so in order to monetize that, we put the walls up. We made it much harder to try to get data from the platform so we could monetize it through the API. I think it also applies in the Wikipedia case, which is we’re running a nonprofit project that a lot of people volunteer contribute to. If we can’t pay for our server costs to make that sustainable, we need to raise the barriers. The promise of the open web originally was that all this knowledge would be free and open. And I guess, Gabe, maybe that’s your cue. It does kind of feel like if we don’t get what Kate is talking about right fast enough, you’ll end up with a web where everybody has pulled up the drawbridge basically.

Gabe Goodhart: Yeah, a hundred percent. And I’m glad you brought up the early web, free and open concept here, because to me there really are two tracks that could happen here: pull up the walls, or build the team and build the partnerships. And I think, in much the same way that open-source software works, where you have large, important projects managed external to any given company, but individual companies which benefit a lot from those projects invest heavily in the maintenance and creation of those projects, the same should be true for high-quality data sources like Wikipedia. Sure, if you’re a grad student writing a scraper, you’re probably not going to also stand up a server and host a mirror of Wikipedia. But if you’re IBM, if you’re Meta, if you’re OpenAI, absolutely that’s a great opportunity to be a positive player in the open data market, and if you could host a portion of the traffic yourself and expose that, now you’ve just built out the ecosystem. And I know there’s the great divide between open and closed in the AI world, but I really think, especially those of us that work at companies who are really leaning into the open side, it’s a great opportunity to actually play well in the space and help lift all ships here. So I would love to see companies like IBM and others partner with Wikipedia to solve this problem at scale rather than necessarily having to bring up the walls.

Tim Hwang: Yeah, for sure. Marina, maybe turning to you, I know earlier you’re like, “Ah, it’s all capitalism,” but this is almost a shift in the social contract of the internet in some ways. The Google era was, well, you make your website open, you let us index it and scrape it, and we promise you that we’ll send you traffic that you can sell ads against. And so you get money for being open in some ways. But AI has less of that feature, right? ‘Cause you build the model and then there’s no return traffic to the sources. And so it almost assumes what Gabe is talking about, I guess, is that the leading companies ultimately have to directly transfer money in some ways to these projects. But curious to add, how you think about it. It feels like in some ways AI is proposing a very different way of how value gets exchanged on the internet.

Marina Danilevsky: Yes and no. So Wikipedia has been in demand for decades. Grad students have been writing scrapers—it’s a long tradition of bad Python scripts that no one reuses. The difference is that when we used to do it before, you didn’t need that much data. You were doing things with topic modeling, you were doing things with graphs and things of that nature. You didn’t actually need all of Wikipedia. If you’re doing things with large language models, yeah, you kind of need all of Wikipedia, and in many ways it’s easier to write a scraper than to go and hunt around and transform data. If you write the scraper, you set it, you go away, and it’s been scraped, and that ends up being a problem. So that hasn’t changed. But right now I think it’s the scale that’s provided the problem, not Wikipedia itself. That’s why it’s the infrastructure, not the knowledge, that’s really provided the problem with AI models. Yeah, I agree that they have consumed the knowledge and they’re not necessarily helping anyone. But I will say that most of the time, when you get an answer from some AI model, you then wanna take an action. The next step is to your benefit to be able to send people back to sources that are trusted, just like the way that right now you search on Google or whatever, it’s gonna send you links that you’re still probably gonna eventually wanna follow. Just like with social media, people are going to learn how to interact with this technology, and they’re going to learn to not just take the AI overview. And even though short-term you might say, “Oh, who cares where this came from?” longer-term we’re gonna swing right back to, “No, I care where this came from.” So I want to be able to have that trust, have that trace. So I think that actually if we start that work now, it’s really gonna be helping ourselves in the future to set that up and have that going. Not in the least because also it would be nice to not kill Wikipedia. Please, we don’t have enough. I really depend on that. So again, I think we’ll be able to get past this, but the more we can get, as Gabe was saying, actors from the larger companies that recognize this is actually in their own interests—not just altruism, it really is in their own interest to do this—the faster we get there, the better.

Tim Hwang: Yeah. And maybe in the near term, people are like, “Oh, the AI overview is just fine. I just use it.” And then after a while people are like, “Hmm, I don’t know about that.” And so there’s a dip in traffic, and then it kind of comes back as people are like, “I gotta check the actual page.”

Kate Soule: I also think there’s too much of a premium on recency to have all of this content just get baked into the model and never go back to the page. So something is visiting these pages to pull the most recent content and then feed that into the chatbot and return an answer, which provides that pass-through opportunity to then click on the link, see the full source, and everything else. So I agree with Marina. I don’t think it’s gonna necessarily revolutionize the value exchange, even if we have to continue to evolve a little bit. What I do wonder though, is how can we move to more of a Common Crawl version of the world where these model providers didn’t all crawl the entire internet ourselves independently? We all started from the Common Crawl snapshots of the internet and used that. And I think we do need, just like the community needs to come to these data providers, for data providers like Wikipedia that are prioritizing the public dissemination of knowledge as part of their mission, we do need to work with them to set up more processes and offerings that are designed and tailored for model providers. So unless they’re saying, “Don’t crawl us, here’s a robots.txt, and we’re not interested in this data ever being used by models”—and it doesn’t sound like Wikipedia is saying that—then it would be great to work together to identify an offering that Wikipedia is gonna start to more purposefully put together to reduce crawl traffic and improve the access of their information to large language models for this new mechanism of consumption.

Tim Hwang: Yeah, and that’s why I’m kind of optimistic in some ways. If some of this is, as Marina’s model was, just how easy it is to get the data... if it’s a nicely produced dataset that’s updated and refreshed and great for your use, there’s no reason to write a scraper. And so there’s a lot of need to build these solutions that almost lower the cost of accessing the data without having to hammer the servers all the time. Great. Well, I’ll move us on to our last story, which is mostly just a fun one that popped up across my radar. There was a story about the Beijing Humanoid Robot half marathon, which featured 1200 human runners alongside 20 robot teams from private companies and various state-backed projects, where they had a robot running alongside the marathon runners. It ends up the humans are still good at this. The winner was able to complete the half marathon in an hour less than the humanoid robot, which still made it across the finish line, but at two hours, 40 minutes, and 27 seconds. And I wanted to bring this up both ‘cause the video is hilarious and you should check it out, and it’s a fun thing to watch, but also because we have been quite skeptical on this show about the entire craze around humanoid robots. Every time I’ve brought it up as a topic, everybody has been like, “This is never gonna be useful. This is just VC theater. I don’t even know why people are talking about this.” But there’s a part of me that saw this happen and thought, this technology, maybe as novelty as it is, seems to be getting quite good. And so I guess, Gabe, maybe you’re new to the show, so maybe I’ll kick it to you first. Are you similarly like, “This is just a play thing,” or do you feel like you’re a little more bullish on humanoid robots being something that actually ends up being practically useful?

Gabe Goodhart: Humanoid robot or not, I am a little bullish on the idea of really thinking wide about modalities and how AI as a general concept interacts with humans. Going back to what I said earlier, I think the real novelty that was the jump between the pre-gen AI days and the gen AI days was taking the need for complex output programming out of the picture and bringing the AI directly into a space where it could robustly interact with humans. And obviously we’ve done a whole lot since then to blend those two things. Every AI model you hit behind a service is in fact a system and not just a model. But I think the idea of extending that beyond the computer screen and the keyboard is actually really interesting. And I don’t know necessarily whether chasing C-3PO is the right direction for that, but I do think there are a lot of places where the physical interaction of robots is actually... right now, running a half marathon is a very constrained scope. And so part of me really wants to unbox what they actually built and understand, okay, now can that exact same robot also—to lean on an internet trope—fold my laundry for me? Probably not, but it’s really interesting to think about the direction of extending into that physical modality as yet another place where AI can meet humans. So the concrete implementation here, I don’t know, I’d have to read a lot of papers about it to understand whether there’s value there, but the chase of bringing AI closer to where humans interact in more modalities I think is pretty cool.

Tim Hwang: Marina, I feel like we’re really trolling you this episode. Every single story, you’re just shaking your head and muttering under your breath. Do you wanna give your hot take on this?

Marina Danilevsky: So, look, there’s value in VC theater. If it means that VCs are gonna give money for actually valuable and timely work in this direction, then great. The same people who built this robot know a lot about robotics in general, so they’re gonna be doing a whole lot of work. So great, bring on the theater. There’s probably aspects here that are interesting in terms of artificial limbs, in terms of movement in general. I mean, you don’t need this thing to be fast. You want it fast? Go get the MIT cheetah robot; that thing’s gonna run real fast. That’s not the point of this either. So honestly, hooray for theater, and as long as it keeps attention on this and all the directions that Gabe just mentioned, these are things that we should continue to do. Just like basic scientific exploration... I feel like people have, in the way that gen AI is right now and the speed at which things are going, everybody is just like, “Great. So what’s the value? What’s the value? What’s the value?” So now I’m gonna disagree with myself ‘cause this is what I say most of the time: “Where’s the value? Wise? There’s no value.” But sometimes you need to allow people the time and the space and the money to do basic scientific research without really knowing what the hell the value is yet. It will eventually come.

Tim Hwang: Kate, I guess maybe I’ll end this episode with a fun question for you: are there other human-robot competitions that you would wanna see robots competing in? I don’t know if there’s particular use cases where you’re like, “I don’t know about the science, but it would be really funny to watch X, Y, Z.”

Kate Soule: Yeah, I don’t know about that one. Folding laundry is certainly top of my use case list. I don’t know about a competition there. But you know, I’m not interested in robots that can run faster than me, so just for many reasons, not interested. I don’t see a lot of value, and if the goal is to transport things faster, I think there are other modalities, getting to Gabe’s point about diversity of modalities, that I think are gonna be prioritized. So I agree with Marina. Certainly there is value to some of these demonstrations and setting targets that you then try to meet and exceed, but I do really wish that we could find non-humanoid, more fit-for-purpose work on robots and prioritize some of that. I think just like we see with models where smaller, more fit-for-purpose models can drive a lot of value and can be built more efficiently, I think we’re gonna see the same in robotics more broadly. And so I’m not super bullish on general-purpose humanoid robots that can both run a half marathon and actually help me around the house in my day-to-day.

Tim Hwang: Well, that’s a great note to end on. And I, for one, as someone who’s spending a lot of time folding tiny child laundry right now, I actually think that would be an incredible spectator sport. I would be very excited about the humanoid robot laundry folding ESPN 4:00 AM in the morning televised competition. Well, that’s all the time we have for today. Kate, Marina, as always, great having you on the show. You’re a dynamic duo every time you come on; I feel like there are all these comments where I’m like, “Oh yeah, I never really thought about it that way.” And Gabe, welcome to the show for the first time; hopefully we’ll have you on at some point in the future. Thanks to all you listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify, and podcast platforms everywhere, and we will see you next week on Mixture of Experts.

Learn more about AI

What is artificial intelligence (AI)?

Applications and devices equipped with AI can see and identify objects. They can understand and respond to human language. They can learn from new information and experience. But what is AI?

What is fine-tuning?

It has become a fundamental deep learning technique, particularly in the training process of foundation models used for generative AI. But what is fine-tuning and how does it work?

How to build an AI-powered multimodal RAG system with Docling and Granite®?

In this tutorial, you will use IBM’s Docling and open source IBM® Granite® vision, text-based embeddings and generative AI models to create a retrieval augmented generation (RAG) system.

Stay on top of the AI news with our experts

Follow us on Apple Podcasts and Spotify.

  1. Subscribe to our playlist on YouTube