Apparently, we have AGI now
Originally published in AI Pulse, Digital Leaders’ weekly newsletter.
Hi, I have great news.
Apparently, we have AGI now 🥳
OpenAI president Greg Brockman said “Welcome to the AGI era”.
Retirement, here I come 😎. Now I just need to decide which coast to move to and ask AI to invent a suncream strong enough to protect me in the scorching English February sun.
Oh. Hold on. They’ve said this before… pretty much every time there’s been a new model released…
Back to work and hiding in the shade it is.
Last week I suggested a definition for Artificial General Intelligence (AGI) as the moment when building the next layer of capability stops being primarily a human responsibility. The AI identifies gaps, builds tools, preserves knowledge, tests its work and helps design the training that raises its own capability ceiling.
Now Anthropic has released Fable 5.1, designed to work across applications for hours with minimal supervision. OpenAI says AI played a significant role in supervising GPT-6 Astra’s training.
That sounds like AGI is getting closer. I’ve always been bullish about where this is heading, even if I haven’t booked the retirement party yet.
OpenAI has now reported a solution to the Navier-Stokes Millennium Prize problem, a longstanding mathematical problem about how fluids behave. It says the work used an internal model significantly more capable than Astra, and has published the proof. If the result stands up to scrutiny, that’s a substantial reason to expect more from the models coming next.
For all their marketing prowess (or bluster?), these companies do still need help to explain the power of their AI models. They need an equivalent to “horsepower”, which started as a marketing term itself to help sell more engines. Let me throw a few out there.
Einsteins - “this model has the intelligence of 0.8 Einsteins”.
Darwins - “this model’s self-improvement capacity is 2.5 Darwins”.
Distracted Workers - “this model can work autonomously with an efficiency equivalent to 4 Distracted Workers (as measured by the average work completed by people during this summer’s World Cup)”.
Meanwhile, the money keeps flowing. Europe’s Mistral raised €3 billion. London’s Ineffable Intelligence added six cofounders as it pursues AI that learns from experience. I’ve backed that approach for many years (and successfully used it building an AI poker player). But regardless of my bias, different teams trying different approaches seems like our best chance of making real progress.
Nvidia agreed to buy Hugging Face for $12.93 billion. Hugging Face are the home of open-source AI, and so this acquisition is bringing a major platform for sharing open models into the company already supplying so much of the hardware they run on.
It also invested in Mistral and, as we explored last week, is helping arrange financing intended to mobilise over $500 billion of third-party capital for AI infrastructure over time. If there is an AI bubble, Nvidia must be getting out of breath blowing it up…
I can see why people are willing to spend ahead of demand. I expect AI to become vastly more capable, and we will need infrastructure ready for it. Until we get to something like AGI, getting those capabilities to work reliably in an actual organisation can still be painfully slow and expensive. Someone has to connect the systems, sort out the data, check the results and deal with what happens when it goes wrong. There’s still an awful lot of human-shaped bottleneck in there.
Sooner or later though, people have to get enough value from using it to pay for all that infrastructure. Which brings me back to the “Einsteins” discussion. We need a way to describe their capabilities that tells us something useful about the work they’ll actually do. What work can the model finish, how often does it get it right, and how much human effort does it take to get there?
I expect these models to get much better, and I’d plan for that. I’d also want to know what they can reliably do for me today, and what it costs to get there. This week see if you can pick a task to do, write down what a successful result would look like, and give it to one of these fancy new models (count how long it takes you to check and fix its work too)…
If it does it well, great. If not, put a pin in that and try again next model release. Your very own benchmark. I’ve got my own benchmark testing all the stuff I do - “Chris’s Retirement Test”… nothing has passed it yet, but one day… 🏖️