How will AI agents change search? In episode 63 of Mixture of Experts, host Tim Hwang is joined by Aaron Baughman, Chris Hay and Kate Soule. First, Perplexity released their Comet browser, and there are rumors that OpenAI is next. Who will win the AI browser war? Then, we discuss frontier model transparency amid Anthropic’s call for a “targeted transparency framework”. Does this only support large model developers? Later, Cloudflare blocks AI scrapers. Is this a good thing? Finally, we discuss how AI is showing up at Wimbledon this year and enhancing the fan experience via Match Chat. Tune in to today’s episode of Mixture of Experts.
The opinions expressed in this podcast are solely those of the participants and do not necessarily reflect the views of IBM or any other organization or entity.
Chris Hay: I just fundamentally disagree. I don’t believe in closed models. Right. I use them, obviously. But I really think the best way of being transparent is to publish what data was used to train your model and publish your weights. Then we can all have a good look at it and tell you whether it’s a good model or a bad model, and you can improve it. But if it’s locked behind a door, then how do we know?
Tim Hwang: All that and more on today’s Mixture of Experts. I’m Tim Hwang, and welcome to Mixture of Experts.
Tim Hwang: Each week, MoE brings together a crack team of the most brilliant and entertaining researchers, product leaders, and more to distill and chart a path through the ever-more complex landscape of artificial intelligence. Today, I’m joined by Chris Hay, Distinguished Engineer and CTO of Customer Transformation; Kate Soule, Director of Technical Product Management for Granite; and Aaron Baughman, IBM Fellow and Master Inventor. We have a packed episode today. We’re going to talk about model transparency, AI scrapers, and Wimbledon. But first, I really want to talk about browsers.
I guess we’ll start with our little round-the-horn question—a bit of a personal question, if I can ask. Let me know if this is too invasive. Aaron, I’m curious: what web browser do you use on a daily basis?
Aaron Baughman: Well, if I’m in a conversational mood, then I’ll use ChatGPT. So that’s something I enjoy. If I’m in a search kind of mode, I’ll use Chrome.
Tim Hwang: Cool. That’s great. Chris, how about you?
Chris Hay: I’m still on Netscape Navigator. Old school.
Tim Hwang: Very nice. I respect that. And Kate, what do you think?
Kate Soule: I use Chrome on my laptop and Brave on my cellphone.
Tim Hwang: Okay, nice. Yeah, I’m a big Brave fan here too.
The use of the phrase “browser war” is kind of funny in the AI context. I just want to give a little history here. “Browser war” has been a term of art for competition in the technology space for decades. The first browser war was Internet Explorer versus Netscape Navigator in the late ‘90s. It happened again in the 2000s between Internet Explorer, Firefox, and Chrome. I think it’s really interesting that we appear to be on the cusp of an entirely new browser war.
Perplexity, the AI search company, has launched a new AI-driven browser called Comet. There are also rumors that OpenAI is working on one of these AI-driven browsers as well.
Chris, maybe I’ll start with you. I guess the first thing would be good for our listeners to get an intuition: why are these cutting-edge technology companies working on software that’s really old—decades old at this point? Why is it important?
Chris Hay: I mean, it’s super important because that’s the path to revenue. I’d love to explain it in a different way, but whoever controls the pane of glass controls the revenue. If we think about how we do search today, you go to Google—you type in your browser bar, you don’t even go to the website anymore. You search, get your ten blue links, and click through. Really, everybody wants to be the first place you go for information. Perplexity is an amazing site; it comes back with awesome stuff. ChatGPT is the same. But there’s a friction point: when I do a search, I open my browser and have to go to the site. Usually, I’m googling “Perplexity” because I’m too lazy to do anything else. So I put it in the browser bar, click, and go to the site. They don’t want that experience. What they want is when you open the browser, you go directly to their pane of glass. Then when you interact and search, they are coming back with the search result. This is all about control—ultimately, it’s going to come down to ads and revenue. Obviously, Google’s not going to be happy about this, and there’s not a lot they can do because they’ve just had a whole “you gotta sell Chrome” thing. They can’t suddenly go gunslinging. So this is going to be interesting.
Tim Hwang: Interesting. Aaron, a question for you. This is a bit of a change of tactics, I think, from frontier AI model companies or Perplexity—companies that are primarily AI-first. There was a view that originally, why rebuild the browser? Everybody’s just going to adopt chatbots, and what used to be search will go over to chatbots. This seems like the other way around: they’re saying we also have to launch in a browser form factor. In response to that, your answer to the round-the-horn question seems like you’re doing both. The question is whether over time they converge—there won’t be any browsers anymore; it’ll just be chatbots, and what they’re calling a browser is really just a chatbot.
Aaron Baughman: Yeah, I mean, the AI-first browser wars are here. It’s great to watch and be a part of. Instead of internet search, we’re moving towards conversational, context-aware discovery. I saw a quote: these AI-first browsers are “designed for curiosity and built for answers,” instead of just searching and linking data pieces together. The stakes are really high. As Chris pointed out, all these companies—frontier companies and established ones—are jousting for the default entry point to the web. The web itself is changing; there are a lot of agent tech protocols coming out—agent talking to agent, providing answers through conversations.
What I like to do is hybrid searching: a traditional search browser flanked on the side by OpenAI or Perplexity search capability. Yes, I’m still waiting to get access to that Comet browser. If anyone from Perplexity is listening, take me off that waitlist so I can get in there.
Tim Hwang: No pressure. Kate, one dimension I think is interesting: browsers have a long history as privacy defenders. You mentioned Brave—I’m a big fan. For folks who don’t know, it’s an offshoot of Chrome with the explicit purpose of shielding privacy online. The question is: should we be worried about AI companies entering the space? The norm hasn’t been strongly set. Part of the expectation with ChatGPT is they get all your transcripts. How do you think that will evolve? Will browser norms around privacy hold over time?
Kate Soule: Yeah, it’s a really good point, and I completely agree. Ads and revenue are the primary driver, but I think there’s an important driver for OpenAI: data collection. I feel like I say this every time I’m on this show: it’s all about more ways OpenAI can collect relevant data without restrictions from other browser providers limiting what they can track and gather. So I completely agree. This will open new revenue opportunities from ads and also from improving the model, creating differentiation to attract customers—a nice virtuous cycle. I worry tremendously about privacy risks. As you said, I use Brave for the same reason; I’m fairly privacy-conscious. I don’t think the incentives will hold to maintain that level of privacy with these model providers creating their own browser platform.
Tim Hwang: Yeah, that’s right. Not to mention, ads have been central to the discussion. One nice thing about existing chatbot tools is the subscription model, which builds more trust than a world monetized through ads. That changes a lot of the value prop of why these technologies are cool and exciting. Do you think there will continue to be a market for subscriptions, or is that like early search engine subscriptions—a phase?
Kate Soule: I think we’ll continue to see subscriptions where there are specific terms around data use, particularly for enterprise use cases. Companies will want to pay OpenAI and others to ensure sensitive data is handled in a specific way—terms that don’t necessarily apply to the general web browser. I don’t think subscriptions will go away; there’s too much emphasis on data stewardship for enterprise models, monetized through subscriptions.
Tim Hwang: Chris, do you think this leads to increasing inequality online? A class of people who can afford USD 200 a month for the pro version (or USD 1,000), and everyone else gets the free browser with ads and chatbot experience.
Chris Hay: All the time. You’re breaking my heart—telling me my current predicament. The amount going out of my account to AI companies is incredible—rent, utilities... Luckily, my wife doesn’t watch this podcast, so she doesn’t know. But I think there’s a two-class system because it’s not enough to have one subscription; you want the latest and greatest models. With subscription, you get higher limits and access to latest models. These models are hugely expensive to run; you’ve got to pay somehow. I’d like to see a world where compute becomes ubiquitous, like electricity—metered. I’d like to see more of a grid system for compute. We’re well off that, but this is why open source is so important and why having compute on your laptop is important—getting smaller models as powerful as possible so it’s not reliant on those with the biggest GPUs. That’s rich considering what I spend on AI, but I really think compute should be ubiquitous.
Tim Hwang: Yeah, that’s right. If it’s widely available, you’d just use more—subscribing and using open-source models. Aaron, final thought: in the original browser war, IE vs. Netscape; second, IE vs. Firefox vs. Chrome. So far, we’ve got Comet from Perplexity, OpenAI probably. Any predictions on who jumps in next in the AI browser war?
Aaron Baughman: Yeah, there’s always Microsoft’s Edge and Copilot, browser companies like Vivaldi, Arc. Many companies will jump in. The winner of the AI browser wars—there may not be a singular winner; it might be a conglomeration or aggregation. But they will define how we access knowledge, do tasks, and even think online for the next decade. It’s going to be fascinating to watch.
Tim Hwang: We’ll keep an eye on it. I’m going to move us to our next topic. An interesting post came out from Anthropic recently entitled “The Need for Transparency in Frontier AI.” Anthropic says, “We’re a frontier model company; we think it’s important to be responsible. Here’s our proposal for regulation around frontier model transparency.” They list things: transparency regimes should only apply to largest model developers; publish system cards; make secure development frameworks public; protect whistleblowers; transparency standards—a whole regulatory stack.
Kate, maybe I’ll kick it to you. I’ll play cynic: I’m a big, valuable company. Why call for regulation? Doesn’t that raise costs? Why would Anthropic do this?
Kate Soule: I think Anthropic has good motives; they’ve always taken a conservative, responsible approach. I don’t think we can refute that. But I worry the way they frame suggestions—with cutoffs for only frontier/large language model providers—plays into a story Anthropic has perpetuated (and gotten flak for): that only big model providers can build safe AI, that only a few (namely them) can do this responsibly. They focus on making big model providers transparent, but startups/smaller places without much revenue get exempt, allowing innovation. I think that’s the wrong approach. We need transparency to help understand and manage risks. Risks aren’t just based on the model; they’re based on the application. I don’t care if a model was trained by a tiny startup, Anthropic, or IBM; if it’s used for medical decision-making, there’s a set of risks and a framework we should be transparent about—like inspecting an FDA-approved drug label. They’re straying too close to “only big model providers can do this safely.” We should have a conversation about transparency around application-based risks, not provider- or model-based.
Tim Hwang: I’m sympathetic to the little guy—a startup wanting to deploy AI, faced with safety/transparency regulations, making it harder. You’re saying maybe they should just man up. How do you think about that? It seems like trying to lower the burden.
Kate Soule: If you’re a startup in medical diagnostics making patient care recommendations, I’m sorry, I don’t care if you’re little. There should be basic requirements. We don’t give free passes to drug companies because they’re new; they still go through trials. But that doesn’t mean there aren’t plenty of ways for startups in lower-risk areas (like retail) to build expertise without the same requirements. It should be application-based.
Tim Hwang: Super interesting. Chris, what do you think? Kate is proposing an alternative: instead of model size, think about use case.
Chris Hay: If the only thing stopping you from developing a chemical weapon is a closed-system prompt, we’ve got a lot more work to do. More seriously, I understand where it comes from. Anthropic, as Kate said, is one of the most responsible companies; they’re open in papers and do a lot on safety. That’s great. But I fundamentally disagree. I don’t believe in closed models. I use them, but I think the best way to be transparent is to publish what data was used and publish your weights. Then we can all look and tell you if it’s good or bad, and you can improve. If it’s locked behind a door, how do we know? If I bought a soft drink and they said, “Don’t worry, you won’t die; it’s got secret ingredient X,” I’d buy the water instead. I get the point about labeling, but come on—problems with labeling. Transparency ultimately comes from being more open. If everybody is open, we can fix bad problems with a wider set of eyes.
Tim Hwang: Aaron, let me play devil’s advocate for Anthropic: “The problem with Chris’s proposal is it gives dangerous capabilities to all sorts of people, which is more dangerous than our proposal.” Do you buy that? Openness might bring transparency but not necessarily guarantee safety.
Aaron Baughman: This is a complex problem; I don’t think one shoe fits all. One of my biggest issues is selective transparency creating unfair competitive barriers, counter to what Perplexity and Anthropic want. What is a frontier model? Sometimes startups begin them. There are loopholes limiting application to largest developers (e.g., USD 100M revenue or USD 1B annual capital expenditure to be liable). Frontier models often come from startups. So the most powerful AI technologies have a loophole—they don’t need to follow until a big company acquires them or grows large. This is ambitious, requiring laws and comprehensive evaluation methods. It’s a good start. In July 2023, OpenAI, Google, Microsoft formed the Frontier Model Forum for AI safety protocols. Anthropic’s release helps push it forward, but let’s be careful about selective transparency that might be inherent and counter to what’s happening.
Kate Soule: I wonder if we’re letting perfect be the enemy of good. Open weights are more transparent, and I agree it’s the way to develop responsible AI. But OpenAI won’t open up overnight; Anthropic—these companies have too much commercial value. What Anthropic is doing is trying to create near-term wins for basic standards and transparency, a call to action. Again, they’re approaching it wrong. Transparency should be guided by downstream risk, not arbitrary size/revenue. But some transparency is better than none.
Chris Hay: Even if you don’t open weights—and I’m a fan of opening them—if you say, “This is how we trained, this is the data,” for transparency... reality is only certain companies can take that much data and have compute to train. But what goes into the model is important. If 20,000 pages on making a chemical bomb go in, the model learns that. We could discuss what goes in. Right now, they’re not transparent; model safety cards... many companies aren’t transparent because they think it’s secret sauce. True transparency: tell us the data.
Kate Soule: There are degrees of transparency. They’re not defining “transparent” or exactly what to do. There are great frameworks like Stanford’s Transparency Index for model providers. It could involve sharing data—super important. That’s why we share all data sources behind Granite training. But there are also real evaluations on safety, bias, limitations—absent from discussion. Even that degree is missing from some frontier models; they’re calling for it, a reasonable start.
Aaron Baughman: Quick point: we’ve talked about organizations/companies creating technology, but some onus is on consumers/users. When they use a model, they need to agree to terms. If they don’t, guardrails can detect/flag nefarious usage. It’s on both sides—producers and consumers. Anthropic could take a view on that. And I’d encourage them: selective transparency and a loophole for small companies might be dangerous.
Tim Hwang: Kate, practical note: you’re in the trenches. How is Granite thinking about transparency, model cards? It’s different from Anthropic, but hearing your team’s thoughts would benefit.
Kate Soule: Anthropic recognizes safety is evolving and not well understood; they’re not prescribing safety, just talking about transparency. That’s our approach with Granite: the field evolves quickly; the best way to arm customers/users/developers is to be as open as possible while maintaining responsible data stewardship. We have rigorous data governance/review; we’re open about every dataset/source used to train Granite, shared in technical papers. We have a robust safety framework, but safety is evolving; all we can do is be open, share known/unknown risks, mitigations. That’s why we include tools like Granite Guardian alongside models to manage risks, knowing not everything can be baked into model weights. We’re open about distribution: models are openly distributed under Apache 2.0 on Hugging Face, available in products. We have a security incident reporting program; about to start a white-hat hacking/bug bounty program for Granite. We’re working with partners on safety, involving the community not just in using Granite but reporting issues.
Tim Hwang: Great—have you back for the white-hat bounty. Moving to next topic: Cloudflare, known for internet infrastructure (managing traffic spikes, DDoS protection), recently said a problem they’re noticing is traffic/load from AI scrapers across the web for training data. They have a system for blocking crawlers; rather than websites opting in, they’ll block by default. Big deal for companies relying on crawlers for data acquisition. Sparked controversy over data norms. Cloudflare says: with all these websites/publishers we protect, our mission may be to monetize, ensure Cloudflare and websites benefit, scrapers can’t scrape value without consent.
Aaron, simply: is this good for the web? Does it solve AI scrapers?
Aaron Baughman: Fascinating story as it evolves. The internet is evolving—we just talked about AI browsers; we have AI-based internet, new protocols like ACP (Agent Communication Protocol), MCP (Model Context Protocol). These protocols create a new network of agents. Agents could become new websites—you visit an agent, it on-demand creates a website/data. Traffic to agents garners payment vs. traffic to websites. We’re in a paradigm shift. Cloudflare enforcing permission-based model is a struggle between giving content creators more control (long-term, maybe less relevant as we go to AI internet) and being bad for AI model training/use (in-context learning, access to large datasets). Excited to see where this goes. We need to be careful not to restrict. A decentralized way: using MCP, a tool pulling data from private repositories not publicly available; you control content and could charge by visiting your agent. There’s a balance; we’ll quickly get there.
Tim Hwang: Feels like a tragedy of the commons. The internet was open with expectation you wouldn’t drive a semi-truck through the front door. As agents/crawlers become more active, maintaining openness isn’t costless. This is a natural response, but I feel nostalgic—an enclosure on the web.
Kate Soule: I completely agree. A lot of promise/value from the internet, society has gotten from crawling/snapshots for others to use—beyond models, look at search evolution, historical snapshots of society. It’s a common resource for democratization of information; now tragedy of the commons. We’ll see erosion. I understand why: there are bad actors. Should be symbiotic ecosystem sharing knowledge, building products making it accessible, all benefiting. But we get reports of model providers ignoring robots.txt, crawling pirated content for training—bad practices. Combined with commercially attractive opportunities one-sided, value not distributed, content creators say this no longer benefits us; we need fences. Tricky position. We as an industry abused this resource with irresponsible actors. We’ll lose a tremendously valuable resource. Only way out: healthier discourse between content providers, organizations like Common Crawl, model providers—figure out responsible approach with distributed benefits.
Tim Hwang: Chris, I admit feeling creeped out. Cloudflare stands up for the little publisher, but if they accomplish their vision, they control large swaths of the internet—deciding who scrapes, how much money, etc. CEO might say, “What’s your solution, smart guy?” Should we be creeped out? Easy to cheer for, but leads down a path not great either. This bit might get edited.
Chris Hay: Taking the other side: what’s the alternative? It’s hard; you need big players vs. big players for protection. I hope this is transient. The internet today is designed for human consumption/browsers, not AI. Back to Aaron’s point: protocols like MCP are about understanding those protocols. Headless content: most organizations host content in CMSs headlessly—storing as structured data, mashing with templates, rendering HTML. Source content isn’t stored rendered. I think we’ll adaptively render content to bot types; bots like markdown, so maybe they get markdown, new formats appropriate for training. Like HTML designed as markup with controls, you can put in things like “pay for content.” I think this is transient; better content protocols will be designed. Right now, it’s a free-for-all. I think it’s great Cloudflare is standing up—companies with curated content, paying to create it, valuable content being taken without monetization because they’re not getting clicks from Google. I understand monetization, but flip side: Savior padlocks my shop, tells who’s allowed in—I want a little fire there. Done with good intention, but we’ll see.
Tim Hwang: Kate, wild speculation before closing: putting two and two together, we’re about to get a big clash—AI browsers negotiating deals to access content from publishers, Cloudflare creating a shield, race between who controls browser/content access and interstitial apps like Cloudflare getting a piece. Any theory on who wins? Beginning of both, but interesting.
Kate Soule: It’ll be a dynamic period in the browser wars with these factors. Looking at traditional platform battles, it’s about who creates value attracting users, snowball effect where content providers need to get on that browser for clicks, dominoes fall. First-mover advantages: if someone strikes a good deal with OpenAI/Perplexity on browsers for top content, paired with real value creation from agents/models embedded—whoever lands killer use cases around content and models working together, making it compelling everyone has to use it, will set dominoes falling.
Tim Hwang: Interesting. Moving to last topic: Aaron, typecasting you—every time you’re on, we talk sports. It’s Wimbledon; you’ve been there. What are you working on this year?
Aaron Baughman: I’m very proud of our work and team. We took on the grand challenge of creating a real-time, accurate large-scale system for fans to interact with—core consumer-facing. About creating experience for speed and accuracy around two main applied R&D projects. The Championship is on now; go to wimbledon.com or mobile app to see our work—Match Chat and live likelihood-to-win estimations. Most proud: we came together to reimagine following a match with Match Chat, an interactive real-time AI assistant built on top of a line graph of agents working together to answer fan questions. Sophisticated: many edges/nodes where agents work together. Real-time, large-scale—we put in threads with timeouts; if no response within time, stop and go to a different agent responding faster. Agents run in parallel, pulling in context to answer within reasonable time for all question types.
The other part: live likelihood-to-win—as a match progresses, we show the story of who we think will win based on many factors: performance of play (real-time streaming data on ball hits, aces, form, backhand), score, decay boosters/factors. We built a series of equations to help fans understand, predicated on pre-match likelihood-to-win (an SVM series of models, decision tree—traditional predictive modeling). Combined, it’s pretty cool.
Tim Hwang: Question: last year we talked about this; great to check in as it evolves. You’re architecting/reimagining the fan experience for spectators. What do you expect? These tools are different from old-school tennis watching (head pivoting, game over). Are fans adopting these tools? Certain segments more into it?
Aaron Baughman: We have a paper published at ACM KDD in Barcelona—open source; Google it. We’ve been doing this about five years. Fans are adopting it absolutely; they expect this type of work. It’s challenging but fun to innovate. At the beginning, we looked at a broad fan base; we didn’t want to be a toy store for all. We selectively picked personas to focus on and be very good at that. That’s where AI commentary started—pipelines creating narration for AI video highlights for generalists. Now with Match Chat, you can ask virtually any question—handling the extreme fan wanting every detail to someone wanting to eat a strawberry on site. It’s a whole gamut and will continue to evolve.
Tim Hwang: Great. Last question: models are good for prediction; sports fans love prediction (who’s going to win). Is there a worry that if prediction gets really good, it takes the thrill out of sports? We just know who wins before the game?
Kate Soule: I think it raises the stakes. If we get better at predicting and someone beats expectation, that’s even more hype. Interesting consequences on gambling industry—different episode. We’ve seen great examples (Wimbledon) of how AI augments understanding of the game—improving predictions, predicting small parts (where to look, where the ball/puck will go), guiding, educational aspects. Not just “X percent chance to win.”
Aaron Baughman: Funny: we’ve noticed players looking at predictions sometimes change their playing style—chicken or egg. Great power, great responsibility.
Tim Hwang: Chris, wrap us up. A friend years ago got into baseball, said, “There are good numbers and statistics—why I’m excited.” We’ve talked about enhancing tennis fan experience with AI. Could this work the other way? If you’re into AI, will this pull you into sports, opening new markets for spectating?
Chris Hay: Yeah, I think so. Fans get into sports for different reasons; for some, it’s data and prediction, for others strategy. For technically oriented people, it opens up the game, allows seeing it differently. Good for hardcore fans too—to Kate’s point about where the ball’s going, you think more. It’s fascinating. I once scraped all IPL cricket data from India. I’ve only been to one game, but it was fascinating. I pulled all data, remembered Malcolm Gladwell on school cutoff months and size. I ran data against IPL cricketers and found the same cutoffs: school year starts September; nobody born in June plays cricket in India. You think the game is solved—earlier born, bigger, more you win. Then Sachin Tendulkar comes along—his stats/birth date blow everything out. I know nothing about cricket; people will complain. But it got me interested for a couple weeks because of stats/data. Football stats/data are huge. It brings new dimensions, will get more people interested, change the game, increase engagement. Good thing.
Tim Hwang: Kate, Chris, Aaron—one of my favorite panels for MoE. Thank you for coming back. We’ll have this exact panel on in the near future, I’m certain. Thanks, listeners. If you enjoyed, get us on Apple Podcasts, Spotify, podcast platforms everywhere. See you next week on Mixture of Experts.
An artificial intelligence (AI) agent refers to a system or program that is capable of autonomously performing tasks on behalf of a user or another system by designing its workflow and utilizing available tools.
Applications and devices equipped with AI can see and identify objects. They can understand and respond to human language. They can learn from new information and experience. But what is AI?
AI assistants are built by a foundation model (for example, IBM Granite, Meta’s Llama models or OpenAI’s models). Large language models (LLMs) are a subset of foundation models that specialize in text-related tasks.
Listen to engaging discussions with tech leaders. Watch the latest episodes.