What are the most common uses of ChatGPT? In episode 73 of Mixture of Experts, host Tim Hwang is joined by Lauren McHugh, Martin Keen and Aaron Baughman to talk about a new report, How People Use ChatGPT. Next, Anthropic released an updated version of their economic index. Then, another paper, this one coming out of DeepMind on agent economies. How likely is this? Finally, how practical are AI wearables and what does a future with them look like? All that and more on today’s Mixture of Experts.
The opinions expressed in this podcast are solely the views of the participants and do not necessarily reflect the views of IBM or any other organization or entity.
Aaron Baughman: Whenever I hear people’s hypotheses and I read this paper, I ask myself a question: Is the human race officially on autopilot? Because first we use ChatGPT for help, but now we use it for everything.
Tim Hwang: All that and more on today’s Mixture of Experts. I’m Tim Hwang and welcome to Mixture of Experts. Each week, MoE brings together a panel of innovators who are pushing the frontiers of technology to discuss, debate, and analyze our way through the week’s news in artificial intelligence. Today, I’m joined by a great crew. We’ve got Aaron Baughman, IBM Fellow and Master Inventor; Lauren McHugh, Program Director of AI Open Innovation; and joining us for the very first time, Martin Keen, Master Inventor. This is going to be a great episode.
Today, we’re covering a lot of interesting research. We’ll talk about a great paper out of NBER called How People Use ChatGPT, the latest edition of the Anthropic Economic Index, and a paper out of DeepMind on agent economies. We’re also going to cover a pretty interesting set of demos called AlterEgo, and talk about the recent Meta wearable. But first, as always, we’ve got the news and headlines from Aili. So, Aili, over to you.
Aili McConnon: Hey everyone, I’m Aili McConnon, a Tech News Writer for IBM Think. I’m here with a few AI headlines you might have missed this week. Google’s parent company, Alphabet, joined the $3 trillion club. Yes, that’s trillion with a T. Only three other companies besides Alphabet can brag about having a market cap over $3 trillion. According to a new report from the World Trade Organization, AI could boost the value of trading goods and services by nearly 40% by 2040. Are you ready for the animal internet? Scientists from the University of Glasgow have taught dogs and parrots to interact with touchscreens. The parrots learned how to use their tongues to play music, and the dogs used their paws to call their friends. Want to dive deeper into some of these topics? Subscribe to the Think newsletter linked in the show notes. And now, back to the episode.
Tim Hwang: So for the first segment, I really wanted to talk about a really interesting paper that came out of NBER, which is in some ways the gold standard for working papers in the field of economics. It’s a paper entitled How People Use ChatGPT. In many ways, it’s a straightforward paper that breaks down, for the very first time in a professional, academically-grounded way, how people are using ChatGPT. Lauren, we’re fortunate to have you on the show, because I believe you mentioned that David Deming, one of the authors on the paper, was your old professor.
Lauren McHugh: That’s right.
Tim Hwang: So, kicking off, I’ll give you the opening remarks on this one. I’m curious what you thought about the paper and if there’s anything that stood out to you in terms of trends. Was anything unexpected, or did this really confirm your biases in terms of how people are using this technology?
Lauren McHugh: Yeah, I think knowing a little about the research behind this, what I appreciate is what it takes to actually create a taxonomy around how people are using ChatGPT. They classified different kinds of tasks that people use it for. What stood out to me was that the number one task was gathering information, which is essentially search. So, ChatGPT is really “SearchGPT.” And from an economic impact perspective, the point was to understand whether, for now at least, with its main use cases, this is really more of a search engine 2.0 technology versus what was probably the hypothesis going in—that this is a net-new, category-making technology. So, it’s interesting to see how that will evolve.
But I also think looking at the economic impact of Gen AI through a standalone consumer tool like ChatGPT is really limiting. I think the bigger economic impact will be where that technology, via an API, gets embedded into other software we use every day. When I can go to Amazon and search for great toys for a five-year-old, or go to the New York Times website to get the latest update on a bill in the Senate—those are probably going to be the bigger economic impacts than people using Gen AI through a standalone chat interface like ChatGPT.
Tim Hwang: Yeah, and Martin, I’d love to pull you into this discussion. Similar to Lauren, I read it and thought, wait a minute, we’ve been sold this multi-purpose, general-use technology, but it’s just search. Martin, is that a disappointing outcome? Lauren seems to argue it’s still early and other use cases are maturing. But is there a possibility where the main use of AI just turns out to be search?
Martin Keen: Well, that’s certainly what this report indicates—how many people are using it for search. What I found interesting, as someone who works in education, was that the main use case in education—50% of the use for work in education—was for writing. In some respects, that’s not terribly surprising, because we see the internet full of AI-generated content everywhere. But it is surprising because large language models today generally aren’t very good writers. They have a certain writing style, and no matter how much you prompt them, they tend to revert to that same phrasing.
Initially, I saw that and thought, oh my goodness, that’s a scary thought.
Tim Hwang: That is kind of a grim outcome—just using this for writing.
Martin Keen: But when you look closer, actually two-thirds of the things classified under the writing taxonomy weren’t about creating new content, but working on existing content—editing, critiquing, or translating. That makes a lot more sense because, to me, that’s where the power of this comes in. If you give it your own sample of writing—if I’m trying to explain something to a particular audience and need to put it in terms they’ll understand—a large language model is pretty good at acting as a reviewer or an editor. That’s how I use it, so I was comforted to see that a lot of people in education are doing the same.
Tim Hwang: Yeah, it’s the same way I use it—largely as an editor and critique writer. Though it’s interesting that pure text generation is still a big thing. Lauren put out an interesting hypothesis, and I’ve used this example on the show before, but I’ll use it again because I love it: when they first invented the PC, early ads suggested you could use it to store recipes. It took a while for us to come up with the spreadsheet, and then everything changed. Do you think there’s a similar dynamic here? Should we be surprised that the chat interface might only be good for a couple of things, and we’re still waiting for different ways of interacting with this technology to see different results?
Martin Keen: Yeah, and think about what some of those uses might be—like coding, for example. This report showed that only about 4% of messages are about coding, which is much less than you’d think. When you look at benchmarks, they’re very coding-focused. Other things that surprised me were people talking about using this as a therapist—less than 2% of messages were about relationships and personal reflection. Games and role-playing were a tiny percentage, less than half a percent. So, the use cases people initially thought large language models would be used for aren’t necessarily playing out in what people are actually doing with these chatbots.
Tim Hwang: That’s right. Aaron, maybe it’s your turn to jump into this conversation. One of the greatest ironies I’ve always thought about with LLMs is that you have this very left-brain technology generating a very wordy, feelings-based output. I had a similar reaction to Martin reading the paper: have I just been living in a bubble? For the last ten or twenty episodes of MoE, we talk about code generation applications every other week. It’s what a lot of high-profile use cases are excited about—Claude Code, etc. But what we’re finding here is that it’s a total bubble. If you’re interested in mass adoption, coding is not the thing.
So, my question is: should technology companies and foundation model providers be totally changing what they’re focusing on? It sure seems like Anthropic is spending a lot of time on code applications, but it accounts for such a tiny slice of what people are doing.
Aaron Baughman: Yeah, whenever I hear people’s hypotheses and read this paper, I ask myself: Is the human race officially on autopilot? Because first we use ChatGPT for help, but now we use it for everything. This paper shows it’s moved from a niche tool for tech-savvy users to consumer tech, just like the internet and smartphones once did.
What really stood out to me was the number of people using these tools—about 10% of the world’s population, roughly 800 million people. Early adopters were professionals, but that’s flipped: 70% of usage now comes from non-workers. For workers, it’s mostly used for writing, computer programming, and helping knowledge workers find information in knowledge-intensive jobs. For non-workers, AI is being embedded into everyday life—creating images, art, video, multimodal content, and helping with life patterns like rewriting content.
When you put both groups together, seeking information, practical guidance, and writing account for about 80% of all tasks and topics used within these models. I think the future, as alluded to in this paper, is that these aren’t just assistants anymore—they’re agents. In sports, an assistant does tasks you tell them to do, but an agent is constantly in the background working on your behalf. We’re seeing solopreneurs pop up—individuals with a lot of force and foresight using these tools, giving small teams access to expert-level tools.
The last point is that even though access parity is closing, countries with the highest GDP still have more accessibility to these tools. It still doesn’t mean there’s an equal playing field in how to use these tools as agents.
Tim Hwang: Yeah, definitely. I want to get to that because I think it relates to the Anthropic paper we’ll talk about in the second segment. But first, a final question for Lauren on business strategy, related to what Aaron said about the future being agents and use cases looking different. In the near term, if we don’t figure out agents, does Google end up winning this game? A year ago, I would have said Google was out of the game, but they’re catching up quickly. This report made me think: if the majority use case is still search, does that mean the incumbent search company ultimately triumphs? Most people still go to Google for search, so naturally, an AI product around search would go hand-in-hand. How much do you think this weighs in Google’s favor in winning the early innings of the AI game?
Lauren McHugh: I really think the fact that search is the number one task—or “information gathering,” as they call it—is because there’s still a long way to go in bridging the imagination gap of what we could use generative AI for. It’s not because that’s the singular best application for it. From a business strategy perspective, investing in actual work—product strategy and design to figure out what problems, besides search, can be solved with Gen AI—doing market research, user research, prototyping solutions—that work has been surprisingly limited.
There’s now a wave of entrepreneurs taking that forward—I think there were 36 AI unicorns this year alone, companies with over $1 billion valuation. That’s crazy. I don’t think search will remain the most dominant use case. We should focus on all the other things it could do, which might look more like a long tail. But if we invest in using creative intelligence to figure out what those things are, test them, and eventually build and scale them, that makes more sense.
Aaron Baughman: To me, search implies it’s human-driven—a human has to enter a keyword. I hope in the future, data will find you. We all become magnets for data; we won’t have to actively search. Agents will predict what we want to see, search for us, and provide what we’re looking for. Hopefully, Google will be on the forefront of that.
Tim Hwang: It’s a funny kind of future where you wake up and think, “Oh, this is everything I wanted—I didn’t even realize it.” That sounds a lot like today’s social media feed, and I’m not sure it’s necessarily delivering what we should be getting, even though it thinks it’s what will get the most engagement. It’ll be interesting to think of a world where we’re not just sending stuff to maximize engagement, but personalizing it because we think there’s utility in you receiving this information.
I’m going to move us to our second topic, which is pretty related to what we’ve been touching on. Anthropic also released a major review of how people are using AI technologies—the second edition of their Anthropic Economic Index. The basic idea is they have a lot of people using Claude, and they want to get a better sense of how people are adopting and using AI in the field. It’s super interesting.
The main thing I want to focus on, which is new for this edition, is that they’ve expanded their analysis beyond the United States. Again, in the spirit of getting out of our bubble—not everybody uses AI for coding—it’s useful to think about the international scene and how AI is being adopted. Aaron, I’ll kick off with the point you raised: Anthropic finds a relationship between wealthier countries and adoption of Claude, with specific income distribution differences in the data. Should we be worried about an AI gap? Will wealthy countries adopt this technology, get all the benefits, and leave poorer countries behind?
Aaron Baughman: I think it’s important to find the signal in the noise. I don’t think it’s all about wealth or GDP. What I liked was the Anthropic AI Usage Index they introduced, which looked at “usage density”—a normalized measure adjusting for working age, showing that smaller, tech-advanced countries lead in usage per person. Going back to Martin’s point, it’s about utility: What are people actually getting out of using these tools? Is it useful? Is it actionable? How much are they using it?
There’s certainly a correlation in the paper between high income and more usage density, but there are corners where people in not-so-high-income countries are learning to work with AI, getting through it via remote technologies—taking remote classes, watching Mixture of Experts, picking up tools that are very accessible. Businesses are trusting these tools more, spilling over to individual adoption. But AI adoption isn’t uniform, and we need to be careful about widening this AI gap.
Tim Hwang: Yeah, and on this education point, it’s funny because when I used to work in an AI startup, we’d say that LLMs allow natural language conversation, making them the easiest interface to adopt—you don’t have to learn to program or read a handbook; you just talk. But it seems there’s still a learning curve, even with conversation. Has AI turned out to be harder to use effectively than we thought?
Martin Keen: Yeah, when you see people selling hundred-page prompting guides online, it makes you think maybe we’re back to where we started—needing a big manual to talk to a chatbot. Prompting still massively affects the outcome of the model. I thought it was interesting how they lined up countries in the study, and usage closely correlated to GDP: the higher the GDP, the higher the percentage of people using the model and getting utility from it.
But I’m interested in what doesn’t correlate with that straight line. Countries like Singapore and Canada are significantly over-indexed—something like four times the amount of people in Singapore use Claude than you’d expect given its GDP. It begs the question: What utility are these people getting that others aren’t? And in India, while we think there’s so much talk about using tools for code generation, half of all Claude usage there was for coding tasks. So, people in different countries are getting different utility from these chatbots.
Tim Hwang: Lauren, there’s another way to have this conversation: maybe it’s no surprise rich countries adopt Claude more because you can spend $200 a month on it. The better versions of the model are expensive. As someone who works in open innovation and thinks about open source, do you think this map looks different for open source? Is there a huge dark matter where people don’t use Claude because they don’t want to pay, but use open-source alternatives that are free and maybe better? Do you buy that?
Lauren McHugh: I think if you look at developer populations in countries, GDP has much less impact because if you’re already trained as a developer, you have access to all the tools—models, inference engines, validation frameworks, tuning frameworks. But developers are a small percentage of the population in some countries versus others, which is fundamentally an access and economic issue.
This played out similarly with social media 10 or 15 years ago—the same debate about higher adoption in higher-income countries. The stakes were lower then because social media has more entertainment value than productivity and labor value. What’s most important now is what happens next. When I was living in East Africa, Facebook created Internet.org, making tools available for free by working with telcos and internet providers. The reception was mixed: some were excited, but in India, it was banned within a year because it was better to have no free internet if it’s curated by someone else, potentially creating information monopolies.
So, this report helps bring awareness that adoption and access aren’t equally distributed. What happens next needs to account for how to do this with dignity for the populations meant to be served, not using charity as a cover for market dominance in emerging markets. That’s the most important thing now that this access gap has been identified.
Aaron Baughman: This paper left me with some hope at the end. Geography isn’t necessarily destiny, but where you live affects what you can do. If you live in a low-usage region, you could learn remotely, create your own AI exposure, or find niche industries that adopt AI. There are ways to close the gap—it could be a grassroots movement to increase usage density in geographies that seem hopeless.
Tim Hwang: I’m going to move on to our third topic: an interesting paper from DeepMind called Agent Economies. It’s a bit speculative, but we can debate how speculative. The paper looks at the idea that in the future, we’ll have AI agents in the economy and in many domains, interacting with one another. It points out that will be weird and new, and we’ll need to figure out what to do with it, introducing new risks.
This connects to something we talked about last episode: in job hiring, people use LLMs to submit applications, and HR teams use AI to filter through them—a little algorithmic war that hasn’t been great for anyone. So, Martin, should we be worried about agent economies? The minute we have automation on both sides of a market, things can get out of control.
Martin Keen: We see this in education too: someone writes an article with AI, someone else uses an LLM to summarize it, then uses an LLM to create quiz questions based on the summary. It’s such an interesting thought—how powerful a particular agent can be, and what happens if you connect it to another equally powerful agent. How does that communication work? From a plumbing point of view, I’ve been looking at integrating agents using things like the Agent-to-Agent protocol (A2A), an open-source thing now part of the Linux Foundation, originally from Google. It lets agents talk to each other, discover each other, and so on.
The analogy I heard is making a Lego brick out of an agent. In my IT career, we’ve been there with SOA, CORBA, microservices—here we are again. But this is a whole new scale: not a microservice that writes a file to a database, but agents communicating, asking each other to do things. How do we know what it does is what we asked? Meaning could get lost as it goes down the line. It’s really interesting to think how this agent economy would work, what the first use cases would be, and all the potential unintended consequences.
Tim Hwang: Aaron, are you hopeful? The paper ends on a positive note, suggesting we need to figure out how to make these economies steerable. My skepticism is that one of the most automated markets is financial markets, with algorithmic trading. The stock market has proven hard to steer—when in crisis, we hit a circuit breaker and stop the market. Are steerable markets a promising frame, or will we just turn it off and on again, hoping the system keeps working?
Aaron Baughman: I’m waiting for agents to unionize, demand profit, and even nap breaks. This reminds me of a field I studied in college: evolutionary computing and artificial life—the study of how manmade systems exhibit behaviors characteristic of living systems. These agents are becoming similar, but the enabler now is AI, which is more top-down, logic-driven.
To your question about whether agents can solve problems in stock markets: you don’t have to be the most powerful agent, just the most needed one. It comes down to setting the right distribution of fairness, credit assignment, and incentives for an agent to self-evolve to better solve problems. You’ll still see traditional machine learning embedded in these agents alongside Gen AI, intertwining—LLMs can use outputs of decision trees or support vector machines, and vice versa. That combination will create scalable coordination among agents.
Tim Hwang: Ultimately, that sounds promising. Lauren, last word before we move to our final topic: Agent economies—are you optimistic?
Lauren McHugh: I’m optimistic about what I see as the next phase: agent companies—companies of agents before we get to economies. I look to open-source communities to see if that’s realistic. Two projects are jaw-dropping: MetaGPT, which claims to be a software startup of agents—agents for market research, competitive landscapes, defining requirements, and an engineering team of agents for code generation and deployment. It’s an end-to-end software company.
The other is AI Scientist—a team of agents that can do its own scientific experiment, come up with a publication, and try to get it published. It’s a popular project with about 10,000 stars. You give it a prompt like, “What’s a more efficient way to use LLMs?” It comes up with a hypothesis, designs the experiment, creates or gets data, runs experiments (usually about LLMs itself, so it’s meta). They got one AI-generated paper accepted into ICLR—they worked with the organizing committee, gave a heads-up, and one of three papers was accepted.
Agent economies—I can’t quite wrap my mind around it. First, let’s make agents work better, then agent companies, which would create agent economies.
Martin Keen: I’m wondering, is this just a giant echo chamber? It comes up with a hypothesis, its solution, then peer-reviews itself—is it just saying yes to everything?
Tim Hwang: I like Aaron’s hypothesis that we’ll see other phenomena emerge—agents unionizing, AI scientists arguing over credit, AI engineering teams complaining about product teams. We’re about to get there—the future of the AI economy.
Aaron Baughman: AI bickering.
Tim Hwang: Yeah, AI bickering. You’ve seen AI cooperation; get ready for AI bickering. All right, final topic: a tale of two wearables. A demo from a startup called AlterEgo a few weeks back—fascinating. A guy sitting, doing stuff with AI with no visible interface—no glasses, wearable, or pendant; just a little behind-the-ear device. Very impressive as a demo, though they caveated it’s still a prototype. It’s almost like invisible AI—a small, unobtrusive device using computer vision and language models to assist you throughout the day.
Split-screen to yesterday: Meta had a huge event demoing Meta AI, announcing a multi-hundred-dollar wearable with Ray-Ban—next-gen glasses. They had some flops in the live demo, but reviews are positive. This might be the glasses that finally get AI integration right, showing cool stuff like live translation captions while talking to someone, or pulling up notes.
Two interesting visions of the wearable AI future. Aaron, which is more compelling? A transparent screen in your glasses, or a fully invisible audio voice?
Aaron Baughman: There’s no free lunch—it depends on the environment. This reminds me of working in biometrics for about ten years. There are brain-computer interfaces like EEGs measuring brainwaves for authentication, back in 2005. AlterEgo uses EMG, looking at neuromuscular facial and throat activity to infer what you’ll say based on muscle activations without projecting sound. There are non-invasive methods too, like fMRI or transcranial magnetic stimulation.
It’s about ubiquitous computing: what devices do you wear? I liked that the Ray-Ban uses a consumer device people already use—sunglasses—and adds tech. If you pick an object with an affordance and add AI (which is becoming invisible), it becomes powerful. The less physically intrusive, the better. AlterEgo is getting there, but I’d like to see more technical research published—I only found a 2018 paper. How big is the vocabulary library? There are a lot of questions.
Tim Hwang: Martin, this goes to something we talked about earlier: ChatGPT was supposed to be the easiest interface, but there’s a lot to learn. Aaron says theoretically it’s better to think about sending a text and have AI do it, but that’s almost more difficult than glasses with a mouse-like interface.
Martin Keen: Watching the demo, it looked like these guys had to concentrate hard to make this near-telepathic wearable send a message. It’s supposed to only pass intentional thought, but I hope there’s an approval button before it sends. I’m answering your question, but also looking at Lauren, thinking, “What’s that picture frame behind her?” I don’t want that in the message.
It’s a very impressive piece of tech, but there’s a disconnect between engineering (“Let’s create a wearable using brainwave analysis”) and marketing (creating a video to sell it). One use case in the video: it’s too noisy to talk to a friend in the room, so you need a telepathic wearable. Couldn’t you just use your phone to text or write on paper?
When the Apple Watch was released, the idea was to put phone tech into a tiny watch. One feature from the keynote was heart tap-backs—measuring your pulse and sending it to someone else using the heart rate monitor and haptic engine. That wasn’t a problem anyone was trying to solve, and nobody bought a watch for that. It quickly disappeared.
Looking at Meta Ray-Ban use cases: a person looks at a building and asks, “What style of architecture is this?” AI gives an answer. Cool, but do I need to know that 4 or 5 times a day? Finding daily use for these technologies is key. With the watch, it turned out to be fitness tracking and notifications—90% of use. So, how will we actually use these AI-powered wearables?
Aaron Baughman: I thought there was a miss on use cases. A better use case could be helping people with speech impediments or who can’t speak—about 5-10% of the global population, 400 to 800 million people. That’s a big market, and it shows a helpful use case, creating empathy between people and product. I’d like to see how we can help humanity better than just neat tech.
Tim Hwang: Lauren, final word of today’s episode: telepathy—would you pay for it? More generally, your thoughts on this space. From Martin and Aaron, I hear skepticism on both sides—waiting for a good use case. Do wearables have a future with AI, or is it still speculative?
Lauren McHugh: I think these two have different purposes. Glasses are about a more convenient or usable interface—making existing technology easier to interact with. AlterEgo creates a new communication plane. We have vocalized language, body language, facial expressions; this sits in between, letting you say something without the whole room hearing. Do we need a new communication plane? In select circumstances, maybe.
Sometimes it’d be convenient in an in-person meeting or party—like wanting to leave without saying it out loud, or asking Aaron if he’s ready to talk about a case study without distracting everyone. There are use cases, but I’m not sure they’re worth the cost yet.
Like Aaron said, the most compelling use case from MIT’s research was helping people with MS or other dystrophies communicate. That seems truly invaluable, worth all the research. If it extends to social conveniences, that’d be cool too.
Tim Hwang: That’s a great note to end on. That’s all the time we have for today. Lauren, Aaron, always great to have you on the show. Martin, hope to have you back sometime on MoE. Thanks to all our listeners. If you enjoyed what you heard, you can get us on Apple Podcasts, Spotify, and podcast platforms everywhere. We’ll see you next week on Mixture of Experts.
Listen to engaging discussions with tech leaders. Watch the latest episodes.
An artificial intelligence (AI) agent refers to a system or program that is capable of autonomously performing tasks on behalf of a user or another system. It achieves this goal by designing its workflow and employing available tools.
Applications and devices equipped with AI can see and identify objects. They can understand and respond to human language. They can learn from new information and experience. But what is AI?
Developers build AI assistants on top of foundation models—for example, IBM Granite, Meta’s Llama models, or OpenAI’s models. Large language models (LLMs), which specialize in text-related tasks, represent a subset of foundation models.