The paper trail behind every AGI claim in 2026

The paper trail behind every AGI claim in 2026

Contracts, benchmark costs, and SEC filings checked against every 2026 AGI claim, so you know exactly what to trust.

An empty witness stand lit under a spotlight facing a wall of glowing data screens, representing AGI claims put on trial against the evidence

Image Credit: Leonardo AI

In this article

No, artificial general intelligence has not been built. Every system on the market today, including ChatGPT, Gemini, and Claude, is a narrow or general-purpose tool trained for specific kinds of work, not a mind that generalizes the way a person does. The benchmarks built specifically to test that kind of generalization still put frontier models under 1 percent against a human baseline of 100 percent, a gap wide enough that no single press release closes it.

Use it the next time an AGI headline shows up, whether you are comparing companies before a decision, checking a risk disclosure, or just want a definition that matches the current evidence. The checklists below are built to get reused, not read once and forgotten.

Ask ten people to define artificial general intelligence, or AGI, and you likely get ten different answers. Ask the people building it for a living, and the answers spread even further apart. In March 2026, Nvidia CEO Jensen Huang said on Lex Fridman's podcast, "I think we've achieved AGI." He pointed to an AI agent capable of launching a billion-dollar company as his evidence, a claim Forbes covered in detail. That same month, ARC-AGI-3, a benchmark launched by the ARC Prize Foundation specifically to test the kind of reasoning AGI would need, had leading AI systems solving under 1 percent of its tasks. Human testers solved every one of them.

Both facts are on record. That gap, between what companies call AGI in a press release and what researchers can measure in a benchmark, shapes most of the AGI debate this year. If the outcome ever touches your paycheck, it is also worth noting how OpenAI's own definition frames the goal: outperforming humans at "most economically valuable work." That framing sits close to an older, more familiar argument about which jobs actually hold their value, one USA Beam covered in a piece on jobs that now out-earn a four-year degree, and one that connects directly to a companion breakdown of what each level of college degree actually pays. AGI's economic definition and the ordinary question of which skills still pay are closer than they look.

What is artificial general intelligence?

Artificial general intelligence describes a machine that can understand, learn, and apply knowledge across any intellectual task a human can do, not just the narrow task it was trained on. No system has met that bar yet. Every AI product on the market today, including ChatGPT, Gemini, and Claude, is a narrow or general-purpose tool built for specific kinds of work, not a mind that generalizes the way a person does.

An empty chair lit at the head of a dark roundtable with blank nameplates, representing the unresolved definition of artificial general intelligence

Image Credit: Leonardo AI

The most widely cited artificial general intelligence definition comes from OpenAI's 2018 charter, which defines AGI as "highly autonomous systems that outperform humans at most economically valuable work." IBM defines it as a hypothetical stage where an AI system matches or exceeds human cognitive abilities across any task. Google Cloud describes it the same way, as a machine that can perform any intellectual task a human can.

Notice what these definitions leave out. None of them require consciousness, emotional depth, or creativity in the way people usually mean those words. A Google DeepMind research team that studied the term in 2023 found that AI researchers have never landed on one agreed definition of general artificial intelligence, only a cluster of competing ones built around economic output, generality of skill, or how closely a system's behavior resembles a human being.

The disagreement has real consequences. Sam Altman, OpenAI's CEO, told CNBC that AGI is not a super useful term anymore, since almost nobody agrees on what it describes. Microsoft CEO Satya Nadella has gone further, dismissing OpenAI's own economic definition as a form of benchmark gaming, and offering his own marker instead: a world economy growing at 10 percent a year. Define artificial general intelligence however you like. Just expect an argument from the person funding it.

Artificial general intelligence vs artificial intelligence

Artificial intelligence is the umbrella term for any system that performs tasks normally requiring human intelligence, from spam filters to self-driving cars to the app that recommends your next show. Artificial general intelligence is one specific, unrealized goal inside that much larger field.

Researchers usually sort AI into three tiers. Artificial narrow intelligence, also called weak AI, handles one task or a tight cluster of related tasks and cannot transfer that skill anywhere else. It is the only tier that actually exists in deployed products today: Siri, Netflix recommendations, spam filters, and even ChatGPT by some strict definitions all count as narrow AI, since each one operates inside a fixed scope. Artificial general intelligence, sometimes called strong AI, would match human flexibility across virtually any domain, picking up a new skill the way a person does rather than needing a new model trained from scratch. Artificial superintelligence sits above both and would exceed the best human performance in every field at once.

AI and AGI are not the same thing. Every AI tool in daily use today, no matter how capable, still lives inside artificial narrow intelligence. Artificial general intelligence describes the next tier up, and nobody has reached it yet.

Artificial general intelligence vs generative AI

People confuse generative AI and artificial general intelligence constantly, partly because the same companies build both and market them with similar language. They are not the same thing.

Generative AI describes systems trained to produce new text, images, audio, video, or code by predicting likely outputs based on patterns in training data. ChatGPT, Gemini, Claude, Midjourney, and Sora are all generative AI. They are genuinely useful and improving fast, but they work through pattern completion within the domains they were trained on. A practical example of that narrow, task-bound usefulness shows up in personal finance, where USA Beam tested specific ChatGPT prompts for budgeting and money questions. Those prompts work because the task is bounded and the model has seen thousands of similar examples, not because ChatGPT has decided to learn finance on its own initiative. A generative model does not transfer a lesson from cooking to carpentry, not unless a training process specifically built that behavior in.

Artificial general intelligence describes a different, still hypothetical, capability: a system that reasons, learns, and adapts across domains the way a human does, without needing separate training for each new task. Generative AI is a product category that exists and ships updates every few months. Artificial general intelligence is a research target that nobody has hit.

The confusion carries a business logic. Framing a chatbot upgrade as a step toward AGI reads better in a funding round than framing it as a better autocomplete.

Types of artificial general intelligence

Because nobody agrees on a single definition, researchers built frameworks instead. The most cited one comes from a 2023 Google DeepMind paper, later presented at the ICML conference, which sorts both narrow and general AI into five ascending levels based on how a system's performance compares against human percentiles.

LevelNamePerformance barStatus for general AI
1EmergingEqual to or somewhat better than an unskilled humanReached, per the DeepMind team's 2023 assessment, by systems like GPT-4 and early Gemini
2CompetentAt least the 50th percentile of skilled adultsNot yet reached
3ExpertAt least the 90th percentile of skilled adultsNot yet reached
4VirtuosoAt least the 99th percentile of skilled adultsNot yet reached
5SuperhumanOutperforms 100 percent of humansNot yet reached, equal to artificial superintelligence

The framework separates general performance from narrow performance, and that split matters. Narrow AI has already hit the top level in specific tasks. AlphaFold outperforms every human scientist at predicting how a protein folds. Deep Blue beat world chess champion Garry Kasparov in 1997. The chess engine Stockfish now beats any human alive at chess, every time. None of that counts as general superhuman AGI, because none of those systems can do anything outside the one task they were built for.

Critics of the DeepMind framework point out a real weakness: terms like competent, expert, and virtuoso do not have a fixed, agreed meaning, and the percentile cutoffs of 50, 90, and 99 were chosen without a clearly stated reason. The framework gives the field a shared vocabulary. Whether it gives a reliable way to measure progress is still an open argument among researchers.

Why AI labs quietly water down their own models

Every AGI conversation assumes a benchmark score reflects what a model can actually do. It rarely does, because the version tested in public is not the same version a lab trains internally.

Two identical marble busts under one light, one fully lit and one dimmed behind glass, representing the alignment tax between a raw AI model and its public release

Image Credit: Leonardo AI

OpenAI's own 2022 InstructGPT paper, published on arXiv, coined a term for this: the alignment tax. Researchers found that fine-tuning with human feedback, the same process that makes ChatGPT polite and makes it refuse harmful requests, produced measurable performance drops on public NLP datasets including SQuAD, DROP, HellaSwag, and a French-to-French-to-English benchmark. The tradeoff is documented, not speculative, and it is worth reading the paper directly rather than taking anyone's summary of it, including this one.

The tax is not a one-time cost. It repeats with every safety pass a model goes through before release. A base, pretrained model tends to score higher on raw reasoning and knowledge benchmarks than the instruction-tuned, safety-filtered version that ships as a consumer product. Later research, including a 2024 paper on mitigating the alignment tax through model averaging, confirms the same pattern shows up across newer architectures, not just the original 2022 case.

This matters for AGI claims specifically because it means the public rarely sees a lab's raw capability ceiling. What ships is a deliberately constrained version, and what a lab holds internally, sometimes for months before release, is closer to the version without the tax applied. Researchers describe the gap between the two as a capability overhang. Nobody outside a handful of internal safety teams knows how large that overhang currently runs at any given lab, which is exactly the detail an AGI headline never mentions.

Some of the tax can be reduced. OpenAI's own fix, mixing pretraining gradient updates back into the reinforcement learning process, shrank the regression without eliminating it. None of the published techniques remove the tradeoff completely. A model tuned to be safer and more helpful in conversation will, on some measurable slice of tasks, remain less capable than the raw model underneath it.

Super artificial general intelligence

Super artificial general intelligence, more often called artificial superintelligence or ASI, describes a system that would exceed the best human performance in every field at once, not just match it. Philosopher Nick Bostrom laid out the concept formally in his 2014 book, Superintelligence: Paths, Dangers, Strategies, and it remains the reference point most AI researchers still cite.

In the DeepMind Levels of AGI framework, superhuman general performance and artificial superintelligence are the same thing: level 5, the top of the scale, where a system beats 100 percent of humans across a wide range of tasks, including some tasks no human can do at all. Nobody has built anything close to this at a general level. The DeepMind team's own 2023 assessment placed current systems at level 1, Emerging, two full levels below Competent and three below Expert.

Industry leaders talk about ASI more casually than the research supports. OpenAI's Sam Altman wrote in 2024 that superintelligence could arrive within a few thousand days, or roughly by 2034. Demis Hassabis of Google DeepMind has said his own timeline runs a bit longer, landing somewhere shortly after 2030 rather than in the next year or two. The distance between those two views says a lot about how unsettled this part of the field remains.

Example of artificial general intelligence

Here is the direct answer: no working example of artificial general intelligence exists today. Nobody has built full general intelligence yet. What gets called an AGI example in headlines is usually one of two very different things.

The first is superhuman narrow AI, a system that beats humans badly at one task and cannot do anything else. AlphaFold predicting protein structures, Deep Blue winning at chess, and IBM Watson beating Jeopardy champions Ken Jennings and Brad Rutter in February 2011 all fit here. Each is a real, verified example of a machine outperforming the best humans at a job. None of them can hold a conversation, drive a car, or do anything besides the single task they were built for.

The second is what the DeepMind framework calls Emerging AGI: general-purpose systems that handle a wide range of non-physical tasks at a level roughly equal to an unskilled adult. DeepMind researchers placed GPT-4, Google's Gemini, and Meta's Llama in this category back in 2023. Newer systems, including the GPT-5 family and Claude, extend that same category rather than moving past it, since nobody has published evidence of a general system reaching the next level, Competent, at the 50th percentile of skilled adults across a broad task set.

So if someone shows you a chatbot and calls it AGI, ask which definition they are using. There is a real difference between a model that writes a good email and a mind that can learn any human skill the way a person learns it.

Does artificial general intelligence exist right now

No, not by the definition most AI researchers use. Every major lab, including OpenAI and Google DeepMind, agrees that no system today performs at human level across the full range of intellectual tasks. That has not stopped some of the industry's loudest voices from saying otherwise.

Jensen Huang's own definition was narrow by design: an AI counted as AGI if it could generate serious revenue fast, not if it could reason, learn, or plan the way a person can. Financial coverage of the remark pointed out an obvious incentive: declaring AGI already achieved raises the case for buying more of Nvidia's own chips, since Nvidia's hardware trains and runs the systems in question. That incentive matters more than it looks, given how tight AI chip supply has been. A hidden bottleneck in high-bandwidth memory, not raw chip output, has been quietly capping how fast labs can actually train frontier models, a constraint USA Beam broke down in its look at the 2026 AI chip war. A company sitting at the center of that bottleneck has every reason to talk up the finish line.

A tall unstable stack of glowing computer chips with the top chip cracked, representing the AI chip supply bottleneck behind AGI claims

Image Credit: Leonardo AI

The benchmark evidence backs the skeptics. ARC-AGI-2, a test built specifically to measure abstract reasoning without letting a model lean on memorized patterns, had a human panel scoring around 60 percent while frontier models started near zero when the test launched. When the ARC Prize Foundation released an interactive version, ARC-AGI-3, in March 2026, human testers cleared 100 percent of its environments while the strongest purpose-built AI agent managed just 12.58 percent, and general-purpose frontier models from OpenAI, Google, and Anthropic all scored under 1 percent.

There is one notable dissent worth naming. Anthropic president Daniela Amodei told CNBC that AGI is such a funny term, and said that by some definitions of matching human capability, Claude has already cleared that bar in narrow professional terms, pointing to coding, where it now writes code about as well as many of Anthropic's own engineers. She was equally direct about the limits, noting that Claude still cannot do a lot of things humans do without effort. Her point undercuts the idea of a single AGI finish line: AI systems can now beat humans at some tasks while still failing at closely related ones, which makes a yes-or-no milestone hard to define.

What a benchmark score actually costs to produce

Ask what a model scored on a benchmark and you get half the answer. Ask what that score cost to produce, and the AGI conversation looks different.

OpenAI's o3 system, tested against the original ARC-AGI-1 benchmark in December 2024, is the clearest documented case. The ARC Prize Foundation reported that a low-compute configuration of o3 scored 75.7 percent within its $10,000 leaderboard budget, at a cost of roughly 20 dollars per task. A high-compute configuration of the same model, using about 172 times more compute, reached 87.5 percent, a jump documented in detail below.

A small trophy on a pedestal atop an overflowing pile of receipts, representing the real compute cost behind a high AI benchmark score

Image Credit: Leonardo AI

A later academic write-up, RC-AGI-2 benchmark paper, put the estimated cost of that high-compute run near 20,000 dollars per task. Both scores describe the same underlying model. Only the spending changed.

ConfigurationARC-AGI-1 scoreEstimated cost per task
o3, low-compute mode75.7 percentAbout 20 dollars
o3, high-compute mode (172x)87.5 percentEstimated near 20,000 dollars, per the ARC-AGI-2 paper
Human tester baseline100 percent, panel averageAbout 5 dollars, per ARC Prize's own testing

ARC Prize creator Francois Chollet noted that human testers solved the same tasks for around 5 dollars each. That comparison rarely makes it into headlines that report the 87.5 percent figure as evidence AGI has arrived.

This is not a one-off quirk of a single model. It is a structural feature of test-time compute, the technique behind o3 and its successors, where a model searches through many possible solution paths before committing to an answer. More searching produces a better score. More searching also costs more money and more time. A leaderboard score without a disclosed compute budget attached to it says almost nothing about whether that performance is remotely close to economical at production scale, which is the entire premise behind OpenAI's own economic definition of AGI in the first place.

The practical rule: when a benchmark score gets cited as evidence of a capability milestone, check whether the source discloses the compute budget next to it. If it does not, treat the number as a ceiling reached under unlimited spending, not a working capability anyone could deploy at scale today.

The contract clause investors actually watch

Most AGI explainers treat the term as a purely scientific argument. It is also, quietly, a legal one, and the legal version has changed more recently than most coverage reflects.

For years, OpenAI's commercial agreement with Microsoft included a provision tied to AGI. If OpenAI's own board determined the company had built a system meeting its internal AGI threshold, Microsoft's license to OpenAI's technology could lapse. That structure gave OpenAI's board unusual leverage over its most important commercial relationship, and it fed years of speculation every time a new model launched: did this count?

That changed in October 2025, when OpenAI completed a restructuring into a public benefit corporation and signed a new agreement with Microsoft. Microsoft now holds intellectual property rights to OpenAI's models and products through 2032, including whatever comes after AGI. Any future AGI declaration is now verified by an independent panel of experts rather than decided unilaterally by OpenAI's board, and a revenue-sharing arrangement between the two companies stays in place until AGI is reached, whichever comes first between 2032 and that verification.

A sealed legal contract and pen on a glass boardroom table at night, representing the OpenAI Microsoft AGI contract clause

Image Credit: Leonardo AI

This is worth sitting with, because it flips the old incentive on its head. Under the original clause, OpenAI's board had reason to define AGI narrowly, since declaring it too early could unravel its most valuable partnership. Under the new terms, an early AGI declaration mainly triggers an independent review rather than an automatic termination, which lowers the stakes of the announcement itself, though the revenue-sharing and IP terms still make timing far from neutral. When OpenAI's funding and its March 2026 round from Amazon, Nvidia, and SoftBank get discussed, this contract detail rarely comes up. It should, since a large share of Amazon's commitment to that round was reported to be contingent on OpenAI reaching an IPO or the AGI milestone, whichever arrives first.

What OpenAI told investors that the press releases left out

Press releases describe an AGI clause in a sentence or two. Investor-facing documents describe it in risk factor language, and the two rarely match in tone.

In March 2026, CNBC reported on a financial document OpenAI shared with prospective investors ahead of its funding round, a document formatted like a draft IPO prospectus. It listed OpenAI's dependence on Microsoft as a named business risk, alongside its exposure to a possible global chip shortage and ongoing litigation with Elon Musk's xAI. The same document disclosed that OpenAI had reported 13.1 billion dollars in 2025 revenue against roughly 900 million weekly ChatGPT users at the time of filing. A company does not volunteer a dependency as a risk factor unless it is legally advisable to, or unless it is trying to get ahead of a question an investor would otherwise ask directly.

A glossy press release folder beside a highlighted stack of legal filing pages, representing what AI investor risk disclosures reveal about AGI claims

Image Credit: Leonardo AI

The AGI clause tells a similar story once the legal history gets checked instead of the headline. The original 2019 Microsoft agreement gave OpenAI's board sole authority to declare AGI achieved, a determination that required no independent verification and no Microsoft sign-off, and would have voided Microsoft's commercial license the moment it happened. That version of the clause was renegotiated in October 2025. Microsoft now holds intellectual property rights to OpenAI's models through 2032, any future AGI declaration goes through an independent expert panel rather than OpenAI's board alone, and Microsoft gained the right to pursue AGI development independently using OpenAI's methods after that date. Outside researchers who tracked the language change on OpenAI's own site over time reported that the practical effect is a clause far harder to trigger, with Microsoft's downside risk from an early AGI declaration dropping substantially.

None of this makes either company look dishonest. It is a reminder that the AGI conversation runs on two different registers at once: confident, headline-ready language in public, and hedged, liability-aware language in the documents regulators and investors actually read. A reader who wants the more candid version should default to the filings, not the keynote.

Where the claims stop matching the data

Several specific AGI claims sound credible on first read and fall apart once checked against the record. A short comparison makes the pattern clear.

Two different scales in one room, one balanced and lit, one tilted and dim, representing the gap between AGI claims and verified data

Image Credit: Leonardo AI

The claimWhat the record actually shows
A model beating a benchmark means it has that skill generallyBenchmark contamination is hard to rule out once a model's training data postdates the benchmark's public release, which is why researchers keep building new, uncontaminated tests
AGI will arrive as a single announced momentEvery framework in active use, including DeepMind's levels and OpenAI's own charter language, describes AGI as a gradient. Even lab leaders disagree on which point on that gradient counts.
More compute and bigger models guarantee AGI eventuallyThis is an active, unresolved dispute inside the field, not settled science. Meta's former chief AI scientist, Yann LeCun, left the company at the end of 2025 specifically over this disagreement
Strong coding performance is a good proxy for general intelligenceCoding is one of the most benchmark-saturated, well-defined domains in AI, which makes it a weak stand-in for tasks with less training data and less clear success criteria
Voluntary safety testing means every major lab has ruled out risk before releaseThe federal pre-release review program is non-binding, and Meta has not joined it, unlike OpenAI, Anthropic, Google DeepMind, Microsoft, and xAI

The gap between how AGI gets discussed and how it actually gets measured is wider on almost every specific claim than a general hype versus reality framing suggests.

Two benchmarks, two different answers

A model can lead one benchmark leaderboard and score close to zero on another built to test what sounds like the same thing. That is not a data error. It is a difference in what each test actually measures, and almost no consumer coverage explains why.

Knowledge and reasoning benchmarks like MMLU test whether a model has absorbed a large volume of factual and procedural knowledge during training and can apply familiar reasoning patterns to answer multiple-choice questions. Frontier models now score well above 90 percent on MMLU, a level that sounds close to expert human performance. Novel-task benchmarks like the ARC-AGI series test something structurally different: whether a model can figure out a rule it has never seen before from a handful of examples, the way a person encountering an unfamiliar puzzle would. The same frontier models that clear 90 percent on MMLU scored under 1 percent on ARC-AGI-3 when it launched in March 2026.

A beam of light split by a prism landing on two different scoreboards, representing how AI benchmarks like MMLU and ARC-AGI score the same model differently

Image Credit: Leonardo AI

Both scores are accurate. Neither one is the whole picture, and citing only the flattering one is how a narrow win gets described in general-sounding language. This is a live illustration of Goodhart's law as it applies to AI evaluation: once a benchmark becomes a target that labs train toward, either directly or through data contamination, it stops functioning as a clean measure of the underlying skill it was built to test. ARC-AGI-1 lived this exact cycle, climbing from close to 0 percent with GPT-4o in 2024 to near 98 percent for Google's Gemini 3.1 Pro by early 2026, which is precisely why the ARC Prize Foundation kept building harder successors instead of retiring the project.

The rule worth keeping here mirrors the compute rule above. Never accept a single benchmark citation as proof of a general capability. Ask what a second, structurally different benchmark says about the same model, and treat any gap between the two as informative rather than as noise to explain away.

Is artificial general intelligence possible?

Most people working in the field think so, eventually. The disagreement is mostly about method, not destination.

The dominant approach right now, used by OpenAI, Google DeepMind, Anthropic, and xAI, bets that scaling up today's large language models, adding more training data, more compute, and better reasoning techniques, will eventually produce something close to general intelligence. This is the approach behind GPT-5.6, Gemini, Claude, and Grok. It is also an approach with a real physical cost. Training and running these models depends on data centers that draw enormous amounts of electricity and water for cooling, a tension USA Beam examined in a look at the permit-versus-actual-use gap in AI data center water consumption. The strain is not limited to water either. The physical burden AI data centers place on local electric grids, often before nearby utilities have finished planning for the load, is its own quietly growing story, one USA Beam mapped in a piece on the grid nobody built for this much AI demand. Scaling toward AGI is not just a research question. It is also a resource question that communities near those data centers are already living with.

Not everyone agrees this path works. Yann LeCun, Meta's former chief AI scientist, left the company at the end of 2025 specifically over this disagreement. He has argued publicly that large language models cannot reach human-level understanding no matter how much they scale, because they lack a working model of the physical world. His new company, Advanced Machine Intelligence Labs, is betting on a different architecture built around what researchers call world models rather than text prediction.

A smaller group of researchers goes further, and doubts AGI is achievable at all with current computing approaches, though that remains a minority position even among skeptics. The honest summary: nobody has proven artificial general intelligence is possible, and nobody has proven it is not. The industry is spending hundreds of billions of dollars betting on the first outcome.

Why lab benchmarks lag what researchers already know

A detail that rarely makes it into consumer coverage: the benchmark score attached to a public model release is usually already out of date by the time it reaches a headline.

A dim hallway of vault style doors with footprints leading to one lit door, representing internal AI benchmarks hidden from public view

Image Credit: Leonardo AI

Public benchmarks get gamed within months of release because training data absorbs them. ARC-AGI-1 went from difficult to nearly solved within about two years, with Google's Gemini 3.1 Pro reaching close to 98 percent by early 2026. That is not necessarily because models got smarter at the exact skill the benchmark measured. It is often because the benchmark's questions, or ones structurally similar to them, ended up inside later training data.

This is why labs keep private, held-out evaluation sets that never get published. A public benchmark stops reliably measuring capability the moment a model's training cutoff passes the benchmark's release date, since nobody outside the lab can fully rule out contamination. In practice, this means a model's public benchmark score often reflects a snapshot of capability that internal researchers already knew was outdated, sometimes by six months or more, before the public ever saw the number.

The effect shows up clearly in the ARC-AGI series. Each version resists the shortcuts that made the previous one solvable, and each time, frontier scores reset back toward zero before climbing again over the following months. ARC-AGI-2 reset the gap after ARC-AGI-1 saturated. ARC-AGI-3 reset it again in March 2026. That repeating cycle, more than any single reported score, is the most honest signal available for how close artificial general intelligence really is.

For a reader, the practical lesson is to treat any single benchmark score as a lower bound on how outdated it might already be, not as a live reading of what a lab's best internal systems can currently do.

When will artificial general intelligence be achieved

It depends entirely on who gets asked. Independent researchers and company executives give wildly different answers, and the gap between them is one of the more telling facts in the whole debate.

Start with the researchers. AI Impacts surveyed 2,778 published AI researchers in 2023 and found a median estimate of 2047 for a 50 percent chance of high-level machine intelligence, a system that outperforms humans at every task and does it more cheaply. That estimate had jumped forward 13 years from the same survey conducted just one year earlier, in 2022, when the median was 2060. Only 10 percent of researchers expected that milestone by 2027.

Now compare that with the people running AI companies. Anthropic CEO Dario Amodei has said he expects a model that can do everything a human professional can do, across many fields, sometime in 2026 or 2027. Google DeepMind CEO Demis Hassabis said his own estimate runs longer, landing somewhere shortly after 2030, while Google co-founder Sergey Brin has said he expects AGI before 2030. Elon Musk has predicted AGI by the end of 2026, after predicting it for 2025 the year before that.

That is roughly a 20-year spread between the median researcher and the most bullish executives, and the executives all have a direct financial interest in the answer. Worth remembering the next time a single headline number goes around.

The timeline question has four different answers, not one

Four analog clocks glowing in different colors and showing different times, representing the four different AGI timeline categories

Image Credit: Leonardo AI

Every.AGI timeline article, including most of the ones covering the same executives quoted above, treats the question as if it has one correct date. In practice, the honest answer changes hard depending on which specific capability someone means.

For narrow economic tasks, first-draft writing, customer support scripts, and especially coding, something close to AGI-level performance already exists in specific niches. That is why Daniela Amodei's claim about Claude and coding is not wrong, just narrow to a domain that happens to have enormous amounts of clean training data.

For tasks that require modeling the physical world, robotics, causal reasoning about objects, and cause and effect in messy environments, the timeline stretches much further out. This is the exact gap Yann LeCun cites for leaving Meta, and it is the gap ARC-AGI-3 was built to measure, since its environments require an agent to build an internal model of rules nobody explained to it.

For tasks requiring long-horizon planning across many steps or sessions, current systems remain closer to the starting line than the coding comparison suggests. The sub-1 percent ARC-AGI-3 scores from GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro are the clearest evidence for this category specifically.

For regulatory or safety-gated domains- medical diagnosis, legal judgment, financial advice- even a technically capable model faces a slower real-world timeline, because deployment depends on institutions and liability rules, not just raw model capability.

The rule worth keeping: whenever someone cites a single AGI date, check which of these four categories they are actually describing. A date that fits the coding category rarely applies to the other three.

How close are we to artificial general intelligence

Closer than five years ago, and further than the headlines suggest. Both things are true at once, depending on which yardstick gets used.

By the DeepMind Levels of AGI framework, general-purpose AI systems sit at level 1, Emerging, the same level researchers assigned to GPT-4 back in 2023. No general system has been shown to reach level 2, Competent, which requires matching the 50th percentile of skilled adults across a broad set of tasks, not a handful of benchmarks.

Benchmark results tell a similar story, just with a shorter cycle each time. ARC-AGI-1, a test of visual pattern reasoning built by researcher Francois Chollet, went from difficult to effectively solved within a few years, with Google's Gemini 3.1 Pro reaching about 98 percent by early 2026. Researchers responded with a harder version, ARC-AGI-2, built specifically to resist the shortcuts that made the first version solvable. It reset the gap: human testers still cleared roughly 60 percent of it, while frontier scores climbed back toward 77 percent within about a year. The even harder interactive version, ARC-AGI-3, covered above, reset the scoreboard again in March 2026, with frontier models scoring below 1 percent against a full human sweep.

The pattern repeats every time a benchmark gets harder. Language models climb quickly, someone builds a version that resists the shortcuts that got them there, and scores fall back toward zero. That cycle, more than any single number, is the clearest answer available for how close artificial general intelligence really is.

Artificial general intelligence companies

A handful of companies say building artificial general intelligence is their actual mission, not just a product roadmap. Their strategies, and their balance sheets, differ sharply.

Unfinished skyscrapers surrounded by construction cranes at dusk, representing the unfinished race among AI companies toward AGI

Image Credit: Leonardo AI

OpenAI wrote the definition most of the industry still argues about. The company closed a $122 billion round on March 31, 2026, backed by Amazon, Nvidia, and SoftBank, at an $852 billion post-money valuation, and it reported roughly 2 billion dollars in monthly revenue around the same announcement. OpenAI released its GPT-5.6 model family in July 2026, now the default model behind ChatGPT.

Anthropic was founded in 2021 by former OpenAI researchers, including siblings Dario and Daniela Amodei, who left over disagreements about safety and the pace of commercialization. The company builds Claude and crossed a $30 billion annualized revenue run rate in April 2026, ahead of OpenAI's disclosed figure of roughly $25 billion at the time, despite a far smaller consumer user base and a training budget reported at a fraction of OpenAI's. Anthropic also signed a compute agreement with xAI in May 2026 for access to its Colossus data center, a deal that sits alongside its existing compute partnerships with Google Cloud and Broadcom.

Google DeepMind combines Google's research lab with its product teams and publishes much of the academic work that shapes how the rest of the industry talks about AGI, including the Levels of AGI framework. Its Gemini models compete directly with GPT and Claude, and CEO Demis Hassabis, a Nobel laureate for his work on AlphaFold, has said AGI could still be several years away even as progress speeds up around him.

xAI, founded by Elon Musk in 2023 and now a subsidiary of SpaceX, built one of the largest AI compute clusters in the world at its Memphis facility and released Grok 4.5 in July 2026. Musk has repeatedly predicted AGI within the next year, a prediction that keeps sliding forward by roughly a year at a time. His pattern of parallel infrastructure bets is not new. Starlink, his satellite internet venture, followed a similar path of aggressive buildout years ahead of any guaranteed payoff, a story USA Beam tracked through Starlink's 2026 plans and IPO chatter. Both bets rest on the same underlying assumption: that owning the physical infrastructure layer, whether satellites or GPUs, matters more in the long run than any single product announcement, a point echoed in USA Beam's broader guide to the invisible infrastructure behind wireless communication.

Meta AI took the most public detour. After paying roughly 14.3 billion dollars for a 49 percent stake in Scale AI and hiring its CEO, Alexandr Wang, to lead a new superintelligence effort in mid-2025, Meta lost its longtime chief AI scientist, Yann LeCun, who left at the end of the year over strategic disagreements about whether large language models could ever reach general intelligence. Meta is also, as of mid-2026, the only major American AI developer that has not joined the federal government's voluntary pre-release review program, now run through the Commerce Department's Center for AI Standards and Innovation, which OpenAI, Anthropic, Google DeepMind, Microsoft, and xAI have all joined. The same underlying question, how much authority regulators should have to test or pause a frontier system before release, runs through the AI Kill Switch Act, which USA Beam broke down in a piece on the bill's most debated loophole. Whether that debate ends in binding law or stays another voluntary pledge will shape how the next AGI claim actually gets checked, not just announced.

A five-question check for reading the next AGI headline

This is the part meant for readers who already understand the basics above. It is a working checklist, not a summary, built to be reused every time a new capability claim shows up in a feed.

  1. Which framework is the claim actually using: a percentile-based one like DeepMind's levels, or an economic one like OpenAI's charter language? The same model can look very different depending on which lens gets applied.
  2. Was the benchmark cited released before or after the model's training cutoff? A benchmark released before cutoff can leak into training data through web scrapes even without direct memorization, which inflates the score without proving generality.
  3. Is the claim about a narrow, benchmark-saturated task cluster like coding or math, or a genuinely broad task set? Labs sometimes describe a narrow win using general-sounding language.
  4. Does the person making the claim have revenue, valuation, or contract terms that depend on the claim being believed? The OpenAI-Microsoft IP terms through 2032 are the clearest current example of this kind of incentive.
  5. Has the claim been checked against an interactive, adversarially designed benchmark like ARC-AGI-3, or only against benchmarks the model's training data may already have absorbed?

A reader who runs a claim through these five questions can usually tell within a minute whether it deserves a second look or a shrug.

A working framework for tracking AGI exposure across labs

Everything above this section assumes a reader checking one claim at a time. This part is for a different job: building a habit that holds up across many claims over months, without starting from scratch each time a new headline lands. It assumes the basics from earlier sections are already familiar, so treat it as the advanced layer on top of the five-question check.

A dim war room wall with pinned documents, dates, and red string connecting AI company names, representing a framework for tracking AGI claims over time

Image Credit: Leonardo AI

Start with a simple tracking sheet, one row per company, with columns for any AGI-contingent contract language, the trigger condition, the date it was last modified, and the source document. The OpenAI-Microsoft clause belongs in this sheet. So does any future deal that borrows the same structure, including reports that a share of Amazon's funding commitment to OpenAI's March 2026 round was tied to either an IPO or an AGI milestone, whichever arrived first. A sheet like this turns a scattered memory of past headlines into something a reader can actually check against the next one.

Second, treat compute capacity as a leading indicator, not a lagging one. A lab's next capability claim is downstream of its chip supply and data center buildout months earlier, not the other way around. The hidden bottleneck in high-bandwidth memory that has constrained frontier training runs is a useful example of a variable that predicts future benchmark jumps before any press release does.

Third, build a simple contamination check into the sheet: the benchmark's public release date, next to the model's stated training data cutoff. A claim resting on a benchmark that predates the model's cutoff by more than a few months deserves a discount, since the benchmark or close variants of it may already sit inside the training data.

Fourth, weight executive statements by financial exposure rather than by title or confidence. A claim from someone whose equity, contract terms, or funding round depends on the claim being believed carries different evidentiary weight than the same claim from a researcher with no financial stake in the answer, even when both are speaking in good faith. This is not an accusation of dishonesty. It is a basic discipline anyone reading financial disclosures already applies, and AGI claims deserve the same treatment.

Finally, revisit the sheet on a fixed schedule rather than only when a new headline appears. Quarterly earnings calls, annual safety framework updates, and each new ARC Prize benchmark release are predictable enough to check on a calendar rather than waiting to be surprised by them. The goal is not certainty about when AGI arrives. It is a system that keeps working the next time someone claims it already has.

Common questions about artificial general intelligence

Is ChatGPT artificial general intelligence?

No. OpenAI's own products, including the GPT-5.6 family, are generative AI systems trained for specific kinds of tasks. They perform strongly in many areas, but they do not pick up a brand new domain the way a person does, and they still score poorly on tests like ARC-AGI-3 that measure general reasoning.

Is Claude artificial general intelligence?

No, and Anthropic does not claim otherwise. Its own leadership argues Claude beats many professional engineers at coding while still failing at ordinary tasks people handle without thinking, which reads as advanced narrow AI rather than general intelligence.

What is the alignment tax in AI models?

It is the measurable drop in a model's score on certain benchmarks after safety and human-feedback tuning, first documented in OpenAI's 2022 InstructGPT paper. It means the version of a model that ships to the public is usually not the version with the highest raw capability a lab has access to internally.

Why do AI benchmark scores differ so much between tests?

Different benchmarks measure different things. Knowledge-recall tests like MMLU reward absorbed training data, while novel-reasoning tests like ARC-AGI reward figuring out a rule the model has never seen before. A model can score well on one and poorly on the other without any contradiction in the data.

What is the practical difference between AGI and ASI?

Artificial general intelligence would match human-level performance across most tasks. Artificial superintelligence would exceed the best human performance in every field at once. AGI is the finish line researchers are chasing now. ASI is the next race after that one.

Which company is closest to building AGI?

Nobody can answer that with certainty, since no public benchmark measures general intelligence as a finished product. OpenAI, Google DeepMind, and Anthropic each lead on different measures, including revenue, research output, and specific benchmark scores. None of those measures has settled the larger question.

What happens after artificial general intelligence is built?

Nobody knows, which is exactly why AI safety has become its own research field. OpenAI, Anthropic, and Google DeepMind all publish safety frameworks describing how they plan to test and limit systems as capability grows toward AGI and beyond.

Why the disagreement matters

Artificial general intelligence remains a moving argument between researchers who measure progress in benchmark points and executives who measure it in funding rounds.

The most useful habit for reading AGI news is checking which definition a given claim actually uses: an economic one, a percentile one, a benchmark score, or a press release line, and then checking what it cost to produce and who benefits from it being believed. Jensen Huang's billion-dollar-company bar and the ARC Prize's abstract reasoning tests measure completely different things, and both get reported under the same three letters. The habit of branding a product with the biggest possible claim is not unique to AI labs either. Consumer tech has run the same playbook for years, sometimes with a launch that generates more headlines about its opening numbers than its actual specifications, as when USA Beam covered the Trump Mobile T1 Phone and its reported deposit figures. AGI claims sit on the same spectrum, just with far more money attached to the outcome.

Nobody currently building AI disputes that these systems keep getting more capable, year over year. What people dispute, loudly and often for money, is what to call the moment those capabilities finally add up to something general.

USA Beam take

A single lit candle beside a mirror reflection, symbolizing one verified fact against a spiral of unproven AGI claims

Image Credit: Leonardo AI

Strip away the branding and a few facts hold up on their own. The benchmarks built specifically to resist shortcuts, ARC-AGI-2 and ARC-AGI-3, both show frontier systems scoring far below human baselines on genuinely novel tasks, even as those same systems post strong scores on older, more saturated tests. The compute cost behind a headline score, the alignment tax applied before a model ever ships, and the risk-factor language buried in investor filings all point the same direction: the number a lab shows the public is rarely the full number it has internally. And the loudest AGI claims tend to come from people whose funding, valuation, or contract terms benefit from the claim being believed, while researchers with no financial stake in the answer put the median timeline decades further out. Neither fact requires taking a side on whether AGI is possible. They mean a specific claim deserves a specific check, not a shrug and not applause.

This article is for informational purposes only. It reflects publicly reported figures, benchmark results, contract terms, and filings available at the time of writing, and cites its sources throughout. It is not financial, legal, or investment advice, and it is not a prediction of when or whether artificial general intelligence will be built. Contract terms, funding figures, valuations, and safety commitments can change after publication, so verify current terms directly with the companies or filings involved before making any financial or business decision.

Recent Articles from USABeam

Editor's note: All images accompanying this article were created using AI image generation. All data, figures, and case studies in the article itself are drawn from cited public sources.

Kristal Thapa
Written by

Kristal Thapa

Kristal Thapa is the founder and editor-in-chief of USA Beam, covering U.S. and world news, sports, finance, entertainment, and technology with a commitment to verified information, editorial independence, and clear, fact-based reporting.

About the publisher →