
The AGI declaration's receipts measure spend, not generality
Jensen Huang called Astra AGI while citing compute scale. OpenAI's own numbers show intense research use—not independent evidence that capability transfers generally.
On September 6, Nvidia CEO Jensen Huang declared that AGI had arrived with OpenAI's GPT-6 Astra. The receipts he attached were industrial: more than 100,000 Nvidia Grace Blackwell NVLink72 systems used for training, with another 400,000 GPUs coming online.
The same day, OpenAI reported that its research organisation used 3.1 agent-workdays for every human workday by mid-August. Its median researcher consumed more than $600 of daily inference at API prices; the 90th percentile exceeded $7,000.
These are serious numbers. They show how much compute one frontier lab can turn into research activity. They do not establish that one system can transfer competence generally across unfamiliar work. The AGI declaration's strongest evidence measures capacity and consumption, not generality.
Three receipts answer three different questions
Training scale answers whether a lab could build a large system. Internal token spend answers how intensely its staff use that system. Benchmark scores answer how a particular setup performed on specified tasks.
Generality is a different claim. It asks what transfers to tasks, environments and constraints that were not selected by the claimant. A capacity receipt can make broad capability plausible; it cannot perform that test.
The messenger's position matters without invalidating the message. Nvidia sells the systems being counted, so an AGI frame converts a capability launch into a demand story for more compute. That is not proof the declaration is wrong. It is a reason to keep the claimant, commercial interest and evidence class in the same row of the review.
OpenAI published the caveats beside the numbers
OpenAI's research-acceleration report is more careful than the declaration built around it. It calls the measurements preliminary and says overall research progress probably will not keep pace with the reported activity metrics. High-level planning remained a minimal fraction of agent output tokens, and more than half of successful four-to-eight-hour tasks still involved at least one human intervention.
Even the milestone is scoped. OpenAI defines its automated “research intern” as a system carrying out well-defined research tasks under human direction. That may be extremely useful. It is not a published test of open-ended competence across another organisation's work.
OpenAI chief scientist Jakub Pachocki supplied the governance boundary in a separate essay that day. He wrote that no lab had solved alignment and monitoring well enough to continue scaling at maximum speed for much longer. That warning does not settle whether Astra is AGI either. It shows why a capability label cannot quietly double as a readiness decision.
I put the declaration through a claim ledger
Separate the cited receipts before deciding what the headline earns. The capacity row is full. The internal-use row is full. The independent-transfer row is empty. That is not a verdict on whether AGI exists; it is a verdict on what this evidence can carry into a board deck.
Use five rows when the next capability declaration arrives:
- Claim: write the operational prediction behind the label, not the label itself.
- Capacity: record training and serving scale, but do not treat infrastructure as task performance.
- Activity: record tokens, runtime and experiments, then ask which accepted outcomes changed.
- Transfer: look for independent evaluation on unseen tasks and your own representative workload.
- Control: keep intervention, monitoring and failure evidence beside the capability result.
The useful question is not “Do we agree on AGI?” It is “What should this declaration predict in our eval that the previous model could not do?” If nobody can name that result, the label has no purchasing content yet.
A perfect score can belong to the system around the model
A related archive analysis, shows why benchmark receipts need their setup attached: Nvidia's AVO harness reached 100 on ARC-AGI-3's public set while Nvidia itself warned against reading the comparison as a controlled ablation.
The earlier 30-second AndyMental video below compresses that harness lesson. I inspected its full pack and script, complete 30.04-second render, audio stream, three sampled frames and hosted copy on September 11. It is an older benchmark-specific example, not evidence about Astra or AGI, and the argument remains complete without it.
Compute scale tells you the machine is large. Spend tells you people are using it. A benchmark tells you what one evaluated system did. Keep all three—but make generality earn a receipt of its own.
Do not repeat a capability label until you can name the independent result it predicts on your work.


