Skip to content
semiAIfoundry
Where the Line Was AGI’s finish line, 1843–2026

Analysis · 8 September 2026

Where the line was

For 183 years, the finish line for machine intelligence moved whenever machines reached it. This week, the line changed kind: from whether AI can solve problems whose answers we know to whether it can produce knowledge we did not.

The moving goalpost, 1950 to 2026 An illustrative field. A ball representing machine capability climbs an exponential curve from 1950 to 2026. A goalpost stands at the current definition of intelligence. Each time the ball clears it, the post jumps forward: chess, Jeopardy, Go, language benchmarks, ARC puzzles, the Turing test, an IMO gold, interactive reasoning, and finally original mathematical discovery. Ghost outlines mark where the post used to stand. The current post, past 2026, is blurred and labelled with a question mark.

Each time the ball clears the bar, the post moves. The final marker is different: it crosses from solving known tests into producing a proposed answer at the research frontier. The curve is illustrative; the years are not.

Where the Line Was · 1 minute 56 secondsRevised 8 September 2026

The short version

In 1957, a machine that could beat the world chess champion would be intelligent. In 1997 one did, and it became just brute-force search. That summer, Go was declared the real test, perhaps a century away. In 2016 it fell, and it became pattern matching on a board. In 2019, ARC-AGI was built around novel puzzles language models could not touch. On 2 September 2026, a model scored 99.9% on its third version; the benchmark’s creators wrote, reasonably, that saturation would not prove AGI.

Six days later, OpenAI reported something categorically different. An internal multi-agent system had produced an analytical proof and Lean formalization for finite-time singularity in the forced three-dimensional incompressible Navier–Stokes equations, a proposed resolution of a Millennium Prize Problem that had remained open for roughly 90 years. [77]

This is not another score on a test with known answers. It is a claim to new, formally checkable knowledge. It also reveals that the relevant unit of intelligence is no longer obviously one model: OpenAI describes roughly 10,000 concurrent agents, code and internet tools, cross-group synthesis, 2.7 million messages, and about 130 billion output tokens. Capability now appears at the level of an organized system.

The result still has to survive independent mathematical scrutiny. Clay continues to list the problem as unsolved, and its rules require publication in a qualifying outlet, two years of elapsed time, and general acceptance by the global mathematics community before consideration. [78] [79] That delay is not a footnote. It exposes three different clocks: discovery, verification, and institutional acceptance.

The cycle is older than the computer. In 1843, Ada Lovelace wrote that the Analytical Engine “has no pretensions to originate anything.” Turing quoted the sentence in 1950, named it “Lady Lovelace’s Objection,” and answered that machines surprised him frequently. The objection has been reissued after every milestone in the vocabulary of the day: brute force, lookup, pattern matching, autocomplete, training data. The verdict is 183 years old. Only the noun changes.

This page keeps the ledger: who drew each line, when it fell, what people said afterward, and where the line went next. The latest event changes the final question. If a machine can originate a result that no one knew, the argument can no longer stop at whether it passed our test. It must ask whether the result is correct, attributable, reproducible, and accepted.

The tether between Lovelace’s objection and Turing’s answer

Timeline from 1843 to 2026 showing Ada Lovelace’s objection that machines do only what they are told, Alan Turing’s answer that machines surprise him, and both claims recurring in new forms after major AI milestones.
The tether. Lovelace’s objection and Turing’s answer remain paired across the history of machine intelligence, reissued after each milestone in the vocabulary of the day. [1, 51] semiAIfoundry Research · 8 September 2026

Act one, written in advance

Turing’s list of things a machine will never do

The tether runs through the rest of Turing’s 1950 paper. He catalogued the objections he expected to hear, including the Argument from Various Disabilities: a list of things people would insist a machine could never do. He wrote it down seventy-six years ago. It is the ledger below, in advance. Here is how the list stands this week.

Turing’s own reply to most of the list: they are things people had never seen a machine do, so they assumed no machine could. To “enjoy strawberries and cream” he answered that it “would be idiotic” to try — and then noted that the objection is really about whether anything is felt, which is a different question. Keep that one in mind; it is where the story ends up.

The compression

How long each line stood

Bars start the year a test was proposed as the thing machines couldn’t do, and end the year machines did it. The bars get shorter to the right. That shrinking is the takeoff, visible in the goalposts themselves.

40 yrsChess stood as the test, from Herbert Simon’s 1957 prediction to Deep Blue in 1997.
19 yrsGo stood, from “a hundred years” in 1997 to AlphaGo in 2016.
5 yrsARC-AGI-1 stood, from its 2019 launch to o3 in December 2024.
≈6 monthsARC-AGI-3 stood, from its 2026 launch to 99.9% on 2 September 2026.

“Beaten” is dated to the first widely accepted result at or above human level on the test as originally posed. Where that is contested (the Turing test, driving) the entry in the ledger below says so.

Table view

The ledger

Who drew the line, and what they said when it fell

Each entry records the line as it was drawn, the year it fell, the verdict that followed, and where the line moved next. The verdicts are the point: they are almost always right, and they almost always move the post.

The horizon

Always about twenty years away

Plot the year a prediction was made against the year it said human-level AI would arrive. For seven decades the dots hover in the same band: fifteen to twenty-five years out, whoever is talking. Then, after 2022, the dots fall toward the diagonal — the line where “it arrives” and “I said so” are the same year.

Shaded band: the 15–25-year zone that Armstrong and Sotala found when they analysed 95 predictions in 2012. Survey points are the median year respondents gave a 50% chance of “high-level machine intelligence.” Definitions differ from dot to dot — which is rather the point.

Table view

The pattern

Why the line moves

The ledger is real, but not every revision means the same thing. Three mechanisms recur, and a serious account has to keep them separate.

01 · Denial

The achievement is redescribed until it no longer signifies intelligence.

Search, statistics, pattern matching, and next-token prediction can be accurate mechanistic descriptions. They become goalpost movement when the mechanism is used to erase a capability that the original test was meant to reveal.

02 · Discovery

Success exposes what the test never measured.

Chess did not establish common sense. ARC does not establish reliability in an open institution. Some movement is honest measurement learning: a proxy falls, its limits become visible, and the science improves.

03 · Deployment

A demonstrated capability is not yet an accepted outcome.

Real systems must survive uncertainty, transfer, audit, control, and consequence. Raising the standard from a result to dependable use is not denial. It is the difference between a laboratory event and an operating reality.

It’s part of the history of the field of artificial intelligence that every time somebody figured out how to make a computer do something — play good checkers, solve simple but relatively informal problems — there was a chorus of critics to say, “that’s not thinking.”
Pamela McCorduck, Machines Who Think, 1979
The model layer changes what is possible. The acceptance layer determines what becomes real.
semiAIfoundry Research, 2026

Act four · 8 September 2026

The line left the benchmark.

The benchmark era asks whether a system can match a known answer. The discovery era asks whether it can create an answer nobody had, then expose that answer to proof, replay, criticism, and time.

OpenAI’s account compresses the first two clocks into days. The third is deliberately slower. Together they show why “solved” is an event, a process, and an institutional judgment at once.

88 hours

Discovery clock

From launching the agent effort to the reported analytical resolution.

17 hours

Formal verification clock

Additional time reported for Lean formalization and machine checking.

≥ 2 years

Institutional clock

Clay’s minimum elapsed time, alongside qualifying publication and broad mathematical acceptance.

The machine may have crossed the mathematical frontier in four days. The institution will take years to decide whether the crossing counts.

Provenance is part of the result.

OpenAI says a rumor of other Millennium Problem progress prompted the evaluation. It later learned that the parallel result concerned forced Euler, disclosed the overlap, stated that its agents had not seen that work, and recognized its authors’ priority. In machine-originated research, the lineage of prompts, tools, intermediate results, and concurrent work becomes part of the evidence a field must inspect. [77]

OpenAI states that its proof establishes statements C and D in the official formulation and has released both the paper and a Lean formalization. Clay currently lists Navier–Stokes as unsolved. A formal checker can establish that a theorem follows from encoded definitions and assumptions; independent review must still establish that the encoding, assumptions, and correspondence to the prize problem are complete. Announcement · Analytical paper · Lean formalization · Clay rules

The question

Will it ever end?

The line has moved since 1843. Does that make it inevitable that it always will? The evidence points to three endings: the line can keep retreating, the label can dissolve into ordinary technology, or the argument can migrate from capability to the conditions under which capability becomes trusted knowledge and consequential action.

Ending one

Neverthe default

The line is defined by an absence, and absences refill.

“AI is whatever hasn’t been done yet” is not a joke about the field; it is a description of how the word works. A line drawn at the thing machines can’t do has no stopping rule — when the thing is done, the definition, not the machine, is what has to change. Three forces keep the rule in place. The first is psychological: once a mechanism is understood it stops looking like intelligence (McCorduck’s “chorus of critics,” Brooks’s “that’s just a computation”). The second is structural: machine intelligence is jagged — superhuman on one task, sub-human on its neighbour — so there is always a shortfall to point at, and it is usually the sensorimotor one Moravec named in 1988. [52] [58] The third is interest: for a decade the word carried a contract clause worth tens of billions, and it still carries reputations on both sides. Nobody with a stake in a line wants it settled against them.

Precedent: “only humans can…” The other great moving goalpost of the last century is the one between humans and other animals, and it has never stopped moving. In 1960 Jane Goodall watched chimpanzees strip leaves from twigs to fish for termites; Louis Leakey’s telegram in reply: “Now we must redefine tool, redefine man, or accept chimpanzees as humans.” [63] They redefined man. The line then moved, in order, to language, to mirror self-recognition, to episodic memory, to culture, to teaching, to theory of mind — and each fell to a field study. Sixty-six years on, the line still exists and still moves. Nobody expects it to stop; they expect it to keep retreating toward whatever cannot be observed from outside.
Table view

Ending two

Dissolvesthe likeliest

The question stops being asked.

In 1984 Edsger Dijkstra wrote that whether machines can think is “about as relevant as the question of whether Submarines Can Swim.” [59] Nobody argues that a calculator “really” computes, or that a plane “really” flies; those debates ended not with a verdict but with the technology becoming furniture. Narayanan and Kapoor’s 2025 essay “AI as Normal Technology” is the modern form of the argument: the term retires as the thing becomes infrastructure. [60] This ending is already visibly under way — and, tellingly, it is the people who used the word most who are dropping it. Altman calls AGI “a very sloppy term.” Nadella calls self-declared milestones “nonsensical benchmark hacking.” Brockman, in the week he declared the AGI era, recast it as “a mission concept or spiritual concept.” The contract that once defined it no longer has a trigger. And ARC Prize, having built three benchmarks with “AGI” in the name, now writes that saturating them is “not intended as a finish line,” and measures human-normalised efficiency instead of a verdict. [39] [70]

Turing predicted this ending in the same 1950 paper: “at the end of the century the use of words and general educated opinion will have altered so much that one will be able to speak of machines thinking without expecting to be contradicted.” [1] He was early, as he was on the five-minute test. But the sentence describes exactly how this ends: not with a proof, but with the contradiction going quiet.

Ending three

Changes kindhappening now

The line leaves performance and enters acceptance.

Once a system produces a proposed answer at an open research frontier, another benchmark is no longer the obvious next test. The questions become whether the result is correct, whether the path is attributable, whether independent experts can reproduce it, whether the system can do it repeatedly, and whether institutions can act on what it finds.

Navier–Stokes makes the transition visible. The reported discovery took 88 hours. Formalization took 17 more. Institutional consideration cannot begin until the result has been published for at least two years and broadly accepted. [77] [78] The capability clock can move exponentially while the acceptance clock remains deliberately human.

This is not a softer finish line. It is a more consequential one. The object of measurement has changed from what a model can demonstrate to what a human-machine system can make true, inspectable, and usable in the world.

Where the argument goes when capability runs out

Settles by data

Consequences

What it does to work, science and institutions. Marcus’s bet with Brundage is this in miniature: ten real-world tasks, judged at the end of 2027, six of them arguably done already. Whatever the answer, it is an answer. [72]

Settles by data

Reliability and control

Robustness, monitorability, autonomy — the grounds Marcus gave for waiting this week, and the grounds OpenAI itself gave when it labelled Astra’s cyber capability “critical.” These are measurable, and they will be measured.

Cannot settle

Consciousness and moral status

Turing’s strawberries. The 2023 “Consciousness in AI” report found no current system a strong candidate and no test that could confirm one. [61] This is the goalpost that will not move, because it cannot be reached from outside. It is where the three untestable items on Turing’s list already sit.

A verdict, since the page has been giving them

The proxy cycle will not end on its own terms. A line drawn at the last thing machines cannot do has no stopping rule. Some revisions are denial, some are genuine scientific correction, and some reflect the distance between a capability event and an accepted outcome. All three can occur around the same result.

But the cycle does not need to be won to lose its importance. It ends when the operational question outruns the label. Navier–Stokes sharpens that moment: AGI can remain disputed while machines begin producing claims that alter the frontier of human knowledge. The argument then moves to correctness, control, provenance, consequence, and acceptance. Those are harder questions than a benchmark, but they are questions about reality rather than vocabulary.

Inside the takeoff

Why it feels new and old at the same time

An exponential curve has an odd property: it looks the same at every zoom level. Stand anywhere on it and the past looks flat and the future looks like a wall. Pick any other point and you see exactly the same picture. That is why 1997 felt like the year everything changed, and so did 2016, and so did 2023, and so does this month.

New, because each milestone really is a first: no machine had beaten a world champion, won an IMO gold, learned an unfamiliar game faster than a person, or presented a formally encoded proposed resolution to a Millennium Prize Problem. Old, because the reaction still begins with the ledger: isolate the mechanism, narrow the meaning, and identify the residual.

The latest event is different for one reason. Navier–Stokes was not selected because the current model failed it, and its answer was not waiting in a hidden test set. If the proof survives review, the system did not merely reach a human-drawn line. It extended the field on which the line can be drawn.

The label may remain unsettled. The operating reality will not.