Silicon Valley’s pursuit of human-level AI remains elusive, but its power to reshape society is already here.

Image by Janet Turra & Cambridge Diversity Fund / Better Images of AI / CC-BY
June 8, 2026
A group of hungry visitors arrives at a village and asks for food. When the villagers refuse to share what they have, the visitors fill a pot of water and throw in several stones and talk up a tasty soup they plan to cook. Of course, they announce, if they just had an onion, a carrot, and even some chicken, it’d be more delicious. A villager says they have one of the items, another indicates they might have another ingredient, and the story goes on until there are enough ingredients to make a wonderful soup that residents are happy to eat, along with the visitors. Such goes a version of this European folktale that Alison Gopnik, distinguished professor of psychology and a member of the Berkeley Artificial Intelligence Research Group, tells me when I ask her about the promise of an advanced form of AI that will rival or eventually surpass human intelligence. It is known as artificial general intelligence, or AGI.
“My AI version of that [story] is a couple of heads of tech companies go out to the village of computer users and say, ‘you know, we’ve got artificial general intelligence that we just made with next token prediction and transformers and gradient descent,” she says. “And the users say, ‘that sounds impressive.’ And then they [the tech companies] say, ‘yeah, except we need more data. You guys have any data you could give us?’ and the users say, ‘we’ll give you all of the books that we’ve ever written, all the texts we produced, all the images we have.’”
Of course, it doesn’t end there; the tech companies then indicate that the system is still saying “stupid and obnoxious things” and so ask if the humans could aid in reinforcement learning, a process in which people tell machines which responses are good or not, so the models get better, through trial and error, at responding. And the users comply. Next, the designers indicate that the machine is improving, but what’s missing is some prompt engineering—a training during which users figure out what prompts to give the AI system so it generates even better outputs. And the process continues.
“At the end of the day, the users say, ‘God, it’s amazing; there’s this magical intelligence that just came from a couple of algorithms,’” Gopnik says. “But of course, this is partly a debunking point, because it all depends on the fact that people have done all this thinking.”
This, Gopnik explains, is the moral of the stone soup story, which in this case means ending up with a device that combines, agglomerates, and organizes information from lots of lots of different people. It’s what large language models and large image models do.
Since 2022, when large language models were publicly launched, developers began making a bigger promise: These models could lead to systems that would be as intelligent as humans and able to take on many tasks efficiently. With the promise came a tsunami of investments, massive extraction of resources from consumers, and confusion about what these systems are, how they’re defined, and what it even means to reach the goal of artificial general intelligence. As it turns out, while promising they were on the road to AGI, AI developers are continually morphing the definitions and markers they use. An honest look at where the industry stands with regard to achieving artificial general intelligence—however that’s defined—shows a vast and perhaps permanent gulf between machine and human intelligence.
In the fourth quarter of the 2026 Super Bowl, a mysterious ad by ai.com declared that “AGI is coming” and encouraged users to secure their handles for that website.
Popularized in the early 2000s, artificial general intelligence is a concept for systems that differ from artificial narrow intelligence in that they don’t just focus or excel on a specific task, like a digital voice assistant does, but can take knowledge from different areas and apply it as needed, much as humans do. In short, when achieved, artificial general intelligences will be as smart as humans. The problem is that the definition of the very concept of AGI is murky.
“If you look at [Google] DeepMind’s statements, they’ll say a system that has AGI is able to do most cognitive tasks that humans can do,” says Melanie Mitchell, James B. Alley, Jr. Professor at the Santa Fe Institute. “I’m not totally sure what ‘most’ means or what it involves. If you look at OpenAI’s definition of AGI, they talk about a system that’s able to do most economically valuable tasks, which is kind of a strange definition of intelligence. A lot of things that might be intelligent aren’t economically valuable.”
“There’s an article from Google researcher[s] called Artificial General Intelligence is Already Here,” she adds. “So different people have very different definitions of AGI.”
Peter Norvig, a Director of Research at Google Inc. and the co-author of the article Mitchell referenced, explains that the way researchers build AI software now has become general, as opposed to task-specific or narrow. In the past, software was built to do one thing, such as playing chess, and nothing else. But now some language models can perform several tasks. And that’s general in a way that didn’t exist before. Despite this, he objects to necessarily labeling these AI systems as AGI.
“I wish we didn’t even have that word because it’s not specific enough,” Norvig says. “How general do you have to be? How many different tasks do you have to do? How intelligent do you have to be? Is it the average person, the top one percent person, the best person in the world? And depending on what you choose, it could vary by a factor of a billion of how hard it is to achieve. And it could be everything from we’re already here, to it’ll be decades, to it’ll be never. So, it just doesn’t seem that useful because it’s so vague.”
(Still, Norvig and his co-author of the article, Google vice president Blaise Agüera y Arcas, found the term fairly useful. The first sentence of their piece, published in the magazine Noema, reads “Artificial General Intelligence (AGI) means many different things to different people, but the most important parts of it have already been achieved by the current generation of advanced AI large language models such as ChatGPT, Bard, LLaMA and Claude.”)
Throughout the history of the field of artificial intelligence, researchers have been grappling with moving goalposts. Even defining AI itself has had its share of ambiguity. In the 1950s, scientists created programs such as the Logic Theorist that they considered intelligent, because it could solve math and logic problems. Later, they thought when a machine can recognize certain patterns or play a specific game it will be an example of artificial intelligence.
“Chess playing was a really big thing,” says Julian Togelius, professor of computer science and engineering at NYU and head of AI, Nof1, a website that focuses on financial markets as an AI training environment. So, when machines achieved the ability to beat even chess grandmasters, people wondered if that constituted AI. “Well, no, it’s a search algorithm. Then something else is AI, and it’s always been a moving goalpost. The same thing is true for AGI. I think AGI is in some sense even worse as a concept, because the idea is that it’s general. But what does ‘general’ mean?”
Many artificial intelligence enthusiasts base their arguments that machines will reach human levels of intelligence on the economic definition of AGI, under which the machines will be able to perform almost all economically valuable human tasks. Both Mitchell and Togelius find those arguments to embody a reductionist and incomplete view of human nature.
“Most of what we’re doing is not economically relevant,” Togelius says. “What’s the value of living a good life? What’s the value of being a good friend?”
Even the definition of being economically valuable has changed throughout history. Plowing the fields by hand was once economically valuable; now, creating viral TikTok posts could be deemed as such. What tasks, then, will be considered to have economic value in the future?
“Artificial general intelligence means you can do all these tasks, even the new ones that we haven’t thought of yet,” Togelius says. “But this is a very strange argument, because it becomes doubly unknowable. What are those tasks we don’t know because we haven’t invented them yet, because we haven’t automated the other things yet? So, I think AGI is bullshit: It’s convenient bullshit for some cases.”
That convenience and lack of clarity were hard to miss in AI.com’s Super Bowl ad and in the explanation of the company’s mission provided on its site: “ai.com is on a mission to accelerate the arrival of AGI by building a decentralized network of autonomous, self-improving AI agents that perform real-world tasks for the good of humanity.”

Still, researchers and developers are working toward more powerful systems. Large language models, likely the first very public versions of such systems, analyze massive datasets—thanks to the availability of human-produced information in digital format—and learn to statistically predict the probability of the next word in a sequence of text. This means they work by pattern recognition, rather than rational reasoning. So, while they can dazzle users in some ways (such as spitting out the population of Paris in seconds), they also make rudimentary mistakes and struggle with problems that involve multiple steps to solve.
Now, a generation of programs called large reasoning models (LRM) has arrived. These models are specifically trained to generate the machine version of what is akin to a “chain of thought” for humans. (It’s important to note that while anthropomorphic phrases are often used to describe these systems, such terms “exaggerate AI capabilities and performance by attributing human-like traits to systems that do not possess them.”) Meaning, instead of simply returning an answer, they break down and generate intermediate steps to solve a problem.
To do this, some models are trained on human generated traces of reasoning (based on how a person solves a problem) or traces of reasoning from other language models. Others are trained through reinforcement learning, which allows the model to produce better strategies over time.
After OpenAI’s release of o1 in late 2024, most developers—including DeepSeek, Google Gemini, and Anthropic—launched their own reasoning models, some of which perform impressively on tasks like writing computer code.
“We still don’t have a good sense of what they can do, what their limitations are, why they’re better on these benchmarks,” Mitchell says. “There’s been some claims by companies that they’re showing their work, how they solve problems. But there’s other evidence that what they’re generating—when you see their reasoning traces—is not what they’re doing internally.”
And while AI systems have shown progress on several metrics—such as beating humans on benchmarks like mathematics and programming—those tests might not even be the correct methods to evaluate these systems and compare them to human intelligence.
“We designed them to be hard tests for a person to pass, which doesn’t necessarily mean if a machine passes it, the machine is doing the same thing that that person did,” Norvig says. “It could be coming to the right conclusion in a different way that didn’t really demonstrate the same type of intelligence that we wanted it to. So anytime you have a test, take that with a grain of salt.”
Also, although these tests are focused on specific tasks, it’s unclear whether these systems can take the reasoning capabilities used and apply them to other tasks or to generalize those skills.
Last year, Apple researchers indicated in a white paper that “despite sophisticated self-reflection mechanisms learned by reinforcement learning, these models still fail to develop generalizable reasoning capabilities beyond certain complexity thresholds.” Additionally, the researchers found that while reasoning models perform better than large language models at moderately complex problems they “collapse” when problem complexity is high (as do large language models).

Among the many differences between humans and machines is the ability to plan. Once they decide on a goal, humans can plan and follow it to meet their objective. If things go awry in the process, in most cases people can easily make a new plan or alter course. AI agents, which can act on behalf of and perform tasks for users, can take on some tasks, like scheduling meetings, booking restaurant reservations, or crafting emails. Where they have trouble is if they hit a snag that requires replanning.
“Planning—figuring out the steps that it might take to accomplish something—is in general very difficult for these models,” Mitchell says.
Another thing they’re missing, Mitchell explains, is any kind of ability to reflect on their own thinking or reasoning processes. Unlike humans, who have a sense of how sure they are about an answer to a query, these models have no sense of how certain they are of an answer other than correlations they discover in their training data. Their hallucinations, or false outputs, will not come with a caveat.
“They don’t have a connection to what’s true and what’s not true beyond the statistics of what they’ve learned,” Mitchell says. “I might ask you, ‘what’s the capital of Zimbabwe?’ And you’ll say, ‘I think it might be Harare, but I’m not totally sure.’ You’ll have a sense of how uncertain you are. But these models don’t have that kind of ability to reflect on their own uncertainty.”
Naturally, the models also don’t have what humans call common sense, which can only be gained through experience of physically embodying the world. A self-driving car might not know the difference between a physical stop sign and one depicted on a billboard; it might not recognize a vehicle if it’s overturned and not in the orientation that it was trained on. Because in the real world, there will always be things that aren’t in a model’s training data.
“We mostly train these things by the stuff we’ve written down, and we’ve written a lot as a species, and so there’s a lot of information there, but it’s not quite the same as having lived in the world,” Norvig says. “So, they don’t have experience of actually doing these things themselves, particularly more physical things, and we’re starting to see that.”
Having lived in the world provides humans with an internal representation of how the world works, what researchers call a world model. Even young children possess world models, which they can use to make predictions. And the data they receive and process isn’t as structured or organized as what AI models are trained on. Closely related to having world models, humans also have the capability to understand causality, something that current AI models are not very good at. For example, a reasoning model might be able to indicate that if you drop a glass it would break. But…
“If you give them a new problem with some causal structure that they’ve never seen before, like the problems that we give to kids, the systems are very bad at doing that, even the very large systems with lots of data,” Gopnik explains. “So, figuring out the causal structure of the world is another thing that humans, even two- and three-year-olds, are very good at doing with very small amounts of data, and that the big models are not good at doing, even with enormous amounts of data.”

So how could machines gain real world experience? One option is to have robots—which could serve as the physical bodies of these applications—roam the world and, using sensors and cameras, gather data and perform experiments for large language and reasoning models. And while some of that data-gathering is already in the works, researchers point to the tremendous gap between the capabilities of currently available large models and robotics.
“This is the old Moravec’s paradox that it’s easier to get a machine to beat the world’s best chess master than to get one that can actually pick up the chess pieces,” Gopnik says.
There’s a reason for that: Large language models and large vision-language models (which can also analyze images) have access to massive troves of text and images readily available on the internet. According to an editorial in Science Robotics, large-vision-language models were trained on data so massive that it would take about 100,000 years for a single human to sit down and read all the text. But to train robots, experts need data that’s different than what’s on the internet.
The largest of one type of appropriate dataset collected to train robots–which requires “a combination of video inputs with robot motion commands”—compiles information gained when humans remotely control robots to perform tasks. When converted to hours, the time needed for a human to parse the data amounts to just one year, according to an editorial published last August.
“It’s actually a little more than a year now, because [at the time] there was one particular effort called Physical Intelligence that reported 10,000 hours, and that worked out to be about a year’s worth of data,” says the editorial’s author William S. Floyd distinguished professor of engineering at UC Berkeley, Ken Goldberg. “Now, there are many other companies doing this in parallel, [so it’s] quite a bit more than that, about 22 years’ [worth].”
This type of data collection, called teleoperation, is one option for getting around a deficiency in today’s robots: They don’t have the perception, sensory, spatial and other capabilities needed to complete tasks that are easily performed by most humans. For this type of data collection at this rate, researchers say, it would take humans many years to collect enough data to match a ChatGPT-scale dataset.
“For a language model, the input is text and the output is text,” Goldberg explains. “For robots, the input is images, what’s happening in front of the robot, and the output is motor commands to the robot’s joints.”
A robot, Goldberg explains, has many motors. Each arm has six motors, and each hand has 22 motors. “That’s almost 50 motors just for the two arms; and you have to learn how to control those very precisely to get the robot to do something.”
“This is very different than a language model,” he says. “You hope it could work in the same way: You show it a lot of examples of inputs and outputs and it learns, but it’s much harder than learning language models because it’s a higher dimensional problem.”
The other problem is having the right sensors to relay accurate information. How does one explain, or correctly translate into text, all the sensations of, for example, tying shoelaces or cleaning up a dinner table?
“We don’t have sensors for touch that are very accurate and reliable yet,” Goldberg says. “Now we’re collecting data using mostly vision as the input. We have cameras looking over the shoulder of the robots as the robots are doing things, and then we’re recording. But we’re not collecting what the is robot feeling. So, we’re collecting all this data, but if someone discovers a new way of measuring touch, then we have to start over, because we don’t have that data.”
Researchers have also been placing robots in virtual reality environments, which provides experiential data cheaper and faster than teleoperation. This process aids in collecting some types of data but will still leave the machines lacking real-world knowhow. “Robots trained on simulation data can work well in simulation, but they often fail when manipulating physical objects,” Goldberg argues in Science Robotics.
Both collection methods share a limitation in that they don’t generate enough data close to the volume needed for rapid progress. One alternative is to harvest data from commercially deployed robots that are already tasked with performing operations (like sorting in a warehouse). Yet even this falls short of the scale of data available to train large language models.
“Should robotics approach the problem in a similar way to achieve similar levels of knowledge, it needs to encapsulate all human and other animals’ knowledge of perceiving and doing things,” researcher Aude Billard writes in Science Robotics. “This would require embodying those creatures. However, existing interfaces for transferring knowledge from humans to robots are still cumbersome and intrusive and offer only a crude approximation of human sensing, and we are still a long way from embodying birds, fishes, or worms.”
In addition to the challenges of gathering data, there is the problem of retaining all available data. Companies haven’t solved that problem for large language models, not because it’s not possible, but because it’s not currently profitable.
For example, every conversation between a chatbot and a user is an opportunity for the models to gather more data and use that info for retraining. Instead, when a user closes a session, that information is wiped, because of feasibility and privacy concerns surrounding saving the interactions of so many users and utilizing that information to retrain the models.
“It would be too expensive to have each of those millions of people trying to update the model all at once,” Norvig says. “We know how to do that kind of updating, but we don’t know how to do it in a way that’s cost effective over these thousands and thousands of computers and millions of customers. So, figuring out some way to do that better is really important.”

Despite these hurdles, Big Tech has been dangling AGI like a carrot at the end of an expensive stick. In a report for AI Now, journalist Brian Merchant writes that for OpenAI the term AGI “is most often deployed at crucial junctures in the company’s fundraising history, or when it serves the company to remind the media of the stakes of its mission.”
In short, AGI is a tool for attracting investors, and a successful one at that.
Recently, eight researchers joined forces and hired researchers like Norvig to found a company called Recursive Superintelligence, which, according to its website “embraces the logical conclusion: the fastest path to superintelligence will be realized by AI that recursively improves itself, and does so via open-ended algorithms that drive endless innovation.” The startup, which is months old and initially received funding from the likes of Nvidia and venture capital firm Greycroft, “is now valued at more than $4 billion.”
“According to the 2025 AI Index Report by Stanford University, between 2013 and 2024, total global corporate investment in AI reached $1.6 trillion.” Worldwide AI spending is expected to climb to $2.5 trillion this year, “a 44 percent increase over 2025.” To put these figures in context, NASA’s multi-year Apollo program cost $298 billion adjusted for current dollar value.
Combined, Meta, Amazon, Alphabet, and Microsoft are estimated to spend $700 billion on chips, facilities, and related infrastructure.
It’s impossible to know whether the scale of such investment and spending is sustainable given that the promised technology is not even clearly defined, let alone a certainty. For now, however, stocks for tech companies locked into AI’s orbit seem to be buoying the markets.
“If anything fits with a definition of a financial bubble, this is it,” Togelius says. “Obviously, AI is not going away. These things are going to get better at what they are good at. [But] I do not think they’re going to take over the world and turn us all into paperclips.” In other words, the thought experiment that an AGI superintelligence will divert all resources to one specific task, like creating paperclips, at the expense of all else, and lead to some sort of apocalypse is unlikely to become a reality.
The question then becomes how good the models will get in areas they are currently deficient and to what extent they will become proficient in physical space, an area in which progress has been very slow, Togelius explains. “That progress, if it ever comes, won’t be through generative AI as many companies have us believe,” he says.
This line of thinking parallels renowned AI researcher and Turing Award laureate Yann LeCun’s, who has indicated that to have smarter machines, there need to be systems other than large language models. Perhaps, LeCun suggests, the focus on generative AI is currently taking up all the oxygen in the innovating room. “There is this herd effect where everyone in Silicon Valley has to work on the same thing,” LeCun told The New York Times. “It does not leave much room for other approaches that may be much more promising in the long term.”
Other researchers also believe that the emphasis on, and the investment in, making large language models larger and feeding them more data is misguided.
“If you were going to ask the best scientists, ‘if you could invest a lot of money to get something that can do more of the range of things that humans can do,’ they’d be likely to say, ‘you should have robotics; you should have better perceptual systems,’” Gopnik says. “On the other hand, if you say, ‘how can I make sure that my company does better than another company,’ maybe just making bigger LLMs is going to be a good route for that.”
No matter the route, it is important to consider how these models—whether they are generative AI or more advanced systems—will impact society, the environment, geopolitics, the global economy, and distribution of wealth. “It’s generally true that whenever you have software where essentially the marginal cost of production is zero after you’ve made the first copy, that tends to concentrate wealth in the hands of a few,” Norvig says. “We’ve handled that in the past: We went from an economy 100 years ago that was mostly farmers to one where a couple of percent people can be the farmers for everyone. But that unfolded over decades, and now we’re seeing things unfold over months and years, and I don’t know if our society is resilient enough to face those kinds of changes so quickly.”
Even if society proves robust enough to withstand these changes, an important philosophical query that is as old as humanity remains. From Aristotle to poet Saadi Shirazi to neurologist Oliver Sacks, thinkers have delved into what it is that make humans unique and different from other creatures roaming this Earth. Obviously, there is no simple answer, but empathy, awareness, and creativity are traits that many cite when describing our species. And so, what happens if humans begin to lose some of these abilities as a consequence of integrating these systems into their lives?
“It’s actually literally keeping me up at night. I never was a doomer; I never thought they would take over the world. I still don’t believe that,” Togelius says. “I’m worried about a world where it’s normalized that we outsource too much thinking to the machines, [with] very little room for human excellence.”
Maybe the threat then isn’t that humans will make machines more like us, but that machines will destroy the very traits that make us human, that they will impair our ability to reflect, to create, to love.
The Bulletin elevates expert voices above the noise. But as an independent nonprofit organization, our operations depend on the support of readers like you. Help us continue to deliver quality journalism that holds leaders accountable. Your support of our work at any level is important. In return, we promise our coverage will be understandable, influential, vigilant, solution-oriented, and fair-minded. Together we can make a difference.
Keywords: AGI, AI, Anthropic, OpenAI, artificial general intelligence, generative AI, machine intelligence, reasoning models
Topics: The AI Power Trip
I’m pretty disappointed by this piece. There are cherry-picked quotes from skeptics like Mitchell who don’t believe that AI will get to human level anytime soon, if ever. It would have been great to balance them with opinions by some more concerned experts, e.g. the most-cited computer scientest Yoshua Bengio (see for example https://yoshuabengio.org/en/blog/reasoning-through-arguments-against-taking-ai-safety-seriously) or “AI godfather” and Nobel laureate Geoffrey Hinton. I’m following AI development since my Ph.D. on “expert systems” (symbolic AI) in the 1980s and I’m really concerned by the speed of progress we have seen in the past 5 years and the reckless race for “superintelligence”.… Read more »
💯
This article raises the right questions about AGI’s shifting definitions and the real limitations of current models. But it answers a different question than the one that determines whether we are in danger. The existential risk documented by researchers like Stuart Russell, Yoshua Bengio, or Geoffrey Hinton does not come from a system that “thinks like a human.” It comes from a system that optimizes efficiently enough for an objective we don’t control, with sufficient capabilities to resist correction. These two properties don’t require human-level general intelligence necessarily. The article also overlooks recent empirical findings. Anthropic just published data showing… Read more »