The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?
I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
Random humans don't have a significant cross section of human knowledge available in real-time, although many like to pretend they do, especially in internet comments :P Being able to compete with the capabilities of a median human would be an absolutely world changing achievement.
> I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine.
These machines were designed with the express purpose of being condescendingly sycophant to the point they hallucinate just to state "you are absolutely right".
No wonder some people even find these chatbots to be wife material.
Have you? It’s a dumb model that gets a lot wrong. Codex voice is IMO only good for when I cannot dictate into the good models. The delay introduced by letting it work with a good model kills it for me.
What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?
The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.
> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?
Then it’s an expert system.
Stephen Hawking wasn’t very good at folding clothes.
The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.
I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.
Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.
The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.
Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory.
You only have weights (large immutable memory), or context (small mutable memory).
Humans have mutable long-term memory: I can learn a new skill, adapt an old skill to new information, or learn new knowledge today that I couldn't perform/didn't know yesterday. I don't have a training cutoff.
Context engineering is an attempt to paper over this limitation. You can get really far with context engineering and huge models, but you will never get to AGI because there are many tasks where humans' mutable long-term memory outperforms.
For example, a human can invent a new musical instrument and then learn how to play the instrument they just invented. That's inference (inventing an instrument) leading to training (neuroplasticity). Humans have the ability to train our NNs with considerably fewer training samples. Everything that you can do with transformers is in one causal direction: training -> inference.
Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).
Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.
And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.
Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.
I can't draw a pelican. Literally my only point of reference would be AI pelican drawings from the test. Otherwise I wouldn't know how to draw one at all.
I would be able to draw an accurate bicycle, but I'm an outlier on that. Most people could not draw one [1].
I would maybe argue that Einstein was the most LLM-like of great thinkers.
A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.
A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.
That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.
I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?
This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades.
So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute.
Honestly I don't think I have single great article, though some Quanta ones are ok, and the Wikipedia article is okay.
The best thing I can recommend is what I did, which is ask Claude about the significance of (1) quantum computing error correction, and (2) error correction in black hole holography and research convergence between the two.
There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever.
Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.
AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later.
Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition.
It could be a good theoretical physicist. Actually it could be a good experimental physicist as well since senior experimental physicists use grad students for the manual labor.
For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms.
My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment.
If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to.
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
The last version to fail on those questions was GPT 4.5.
Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".
Nooo. I just asked Claude (Sonnet 5 Medium) how many I's are in assassin, and it said 2. Granted, it got several words correct before this. But no, they still aren't great at this letter-counting thing.
Tokens are the most basic input unit of an LLM. But tokens don't generally correspond to words or letters, rather sub-word sequences. So Strawberry might be broken up into two tokens 'straw' and 'berry'. It has trouble distinguishing features that are "sub-token" like specific letter sequences because it doesn't see letter sequences but just the token as a single atomic unit. 'Straw' and 'r' are two tokens but an LLM is entirely blind to the fact that 'straw' has one 'r' in it.
As an analogy, I might ask you to identify the relative activations of each of the three cone types on your retina as I present some solid color image to your eyes. But of course you can't do this, you simply do not have cognitive access to that information. Individual color experiences are your basic vision tokens.
LLMs see tokens, not words spelled out with letters.
Imagine verbally asking someone who has never seen written text the same question: unless they memorized the answer for the specific word you're asking about, they'd have to guess.
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's are not combined into one token.
We don't know what was going on inside the closed source GPT models, but this paper investigated on some of the open-weight models and found it's not due to tokenization: https://arxiv.org/abs/2604.00778
A lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
If it allows you to distinguish readily between human intelligence and computer intelligence, it seems quite relevant to the question of whether computers have achieved something akin to human intelligence.
It may not be useful for anything else, but at least it can say that.
But it's not relevant as a metric to gauge distance to human intelligence. Humans see individual letters, LLMs do not. If I asked you the relative activation of the cones in your retina as I showed you some solid color image, you couldn't do it. You simply do not have cognitive access to that information. But that says nothing about your intelligence.
A more accurate test would be to give it a list of words (or anything represented as a single token) and ask it how many times that token appeared. I'm sure they have no trouble at that task.
Yes, they have some alien failure modes. But that should be expect, they are an alien intelligence. I might be willing to grant that a lack of ability to reflect on its own level of knowledge is a demerit to it being generally intelligent. But then again it is largely an artifact of training. I suspect if there were a guessing penalty during pretraining they would develop or more readily communicate the strength/reliability of their knowledge.
I am not enthusiastic about criteria for human-like intelligence that imply that dyslexic people don't have human-like intelligence.
[EDITED to add:] I actually don't know whether dyslexic people find it difficult to count letters in words, if they have them already written down by someone else. I suspect they find it harder than people who aren't dyslexic. But perhaps "blind people whose spelling is poor" would have been better; I would not want to deny them human-like intelligence either.
I have slight dyslexia. I can't automatically write double consonants all the time.
But counting certain type of letters is very simple algorithmic task especially in written text. Any reasonably intelligent actually thinking thing should come up with algo and then execute it. Which to me sounds like reasonable minimum bar for general intelligence.
> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.
I don't think it's irrelevant but perhaps not in the way you're assuming. When assessing AGI I'm not evaluating counting characters or even the execution of math operators at any scale or speed. As you observe, computer software from Regex to spreadsheets and Mathematica already handle that well. But AGI isn't about what computers can do, it's about whether AIs can do the specific things which, until now, have been uniquely human capabilities. Like understanding nuanced context and then coming up with novel approaches to solve a new kind of problem not relying on any specific prior training or knowledge (the 'G' is for General).
Most definitions of AGI start from a baseline that already assumes easily passing a Turing test and doing anything via text response that a high school graduate could. I ding LLMs not for failing to count but for failing to intuitively understand the nuanced context of a simple class of problem it hasn't seen in its training data. I fully understand that the reason LLMs fail letter counting is that they operate at the token level. They weren't trained on individual letters first, like human 2nd graders.
The only reason recent LLMs get strawberry and blueberry correct now is that they have those words on their pre-training 'cheat sheet'. However, the underlying fundamental weakness in the way LLM intelligence works which leads to this failure mode still hasn't been addressed. Even when the frontier labs add "recognize any sub-token counting question and write a Python script" to the training cheat sheet so LLMs always pass that test... they'll still be unable to recognize a simple class of problem which isn't on their 'cheat sheet'. As long as that's the case, to me, they aren't AGI because they can't fully replicate human-like recognition of novel problem classes. And it's not just about letter-counting. That gap and others like it lead to many other kinds of non-human brittleness in LLM problem solving. Those are the classes of reasoning, intuition and insight that the ARC-AGI series has been trying to queue up as targets. Not to show how bad LLMs are but to help them be great in all these counter-intuitive edge cases
> planes don't flap wings therefore they cannot fly?
This example still misses my point, which isn't related to usefulness or economic value. I concede that LLMs can have greater utility and economic value than humans on many tasks. The point is most definitions of AGI include something like "can fully replicate all the routine daily tasks done by any competent high-school graduate." That's not related to whether LLMs can solve many high-value problems faster and at larger scale than any human. That was also true of ENIAC in 1946.
The fact an airplane can fly faster and farther than any bird is irrelevant to whether an airplane can "fully replicate all the routine daily tasks done by any competent bird." That's the bird equivalent to most AGI definitions. An airplane can't build a nest or recognize the signals encoded in birdsong.
In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". And in this context, airplanes scoring 15,000% more than birds on 'speed' and 'distance' doesn't matter any more than AIs scoring 15,000% more than humans on 'add 10,000 numbers'. We still aren't near AGI because LLMs cannot fully match any high-schooler's ability to independently conceive new approaches to novel problems not in their prior training data.
which gets us closer to philosophical questions which I'm personally not that interested in.
>In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any".
I'm not sure we want a machine that fully succeeds that test.
Planes pass the 'bird replacement' test on the only criteria that matters to us ... flying.
If we wanted nest making planes I think we'd have them by now. Nest making doesn't rate highly on the problems we're looking to solve though.
I don't want a machine that is moody, or depressed or has schizophrenia, which are all pat of the human condition.
We don't need the human "intuition magic dust" to do 99.99999% of useful work.
They're machines designed to do the work we don't want to. That's as "general" as their intelligence needs to be.
I'd prefer if my clothes folding machine did not have an existential crisis.
>I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"?
Do you possess magnetoreception? a stupid pigeon can "see" the earth's magentic field. why are you blind to it? does a lack of magnetoreception make your intelligence any less "general"
no, you're just blind to it because that's just the way it is.
LLMs are blind to character counting because that's the way they are.
It didn't stop ChatGPT from finding the Jacobian Conjecture counterexample.
Human intelligence and machine intelligence are only going to cross over to a certain degree.
same as plane flight and bird flight are only kinda related.
It matters under the lens of AGI. Artificial intelligence that can meet or exceed human intelligence. Sure, it's good at some things but rather limited at others.
Breadth of capabilities matters... and a promotional video is nice and all, but people are throwing this term around like it's a prize they've won, but they've not gotten there yet.
This implies that human employees don’t have insurance. But they do. My company’s cyber insurance for example covers breaches due to employee mistakes. Most companies also have umbrella liability policies. It’s just that for now, AI “employees” need a LOT more coverage.
If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit.
If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.
What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider.
If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.
The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task?
It's not really a dream though? You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks. Another person can do the same thing but also keep a human in the loop. You and the other person may compete in the market for whatever your product or service is, and you both might do very well or one of you might do better than the other because of a whole host of different reasons. Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor, nor does it require them to offer the LLM service with a guarantee that it will never make any mistakes. So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money.
> What do you mean by bearing no real responsibility for its actions?
If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.
It it makes a mistake and deletes your website from AWS, who is responsible?
If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible?
> If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.
Something in your prompt led it to do that, is alex0015's point. The statistical odds of these frontier models screwing up to that extent are so impossibly low that it would almost have to be intentional or accidental negligence on the part of the prompt writer to accidentally have their agent write pornography to their website.
The burden of the mistake would have to fall on the person that gave the tool instructions, because it can't know that what it did was wrong. Wrong is subjective in this case. It only did what it did because you, figuratively speaking, encouraged it to.
Yes but I'd actually go even further than that. We've all had experiences where the model straight-up produces gibberish sometimes, right? It's happened to me with a badly configured harness on a local model, and even also on frontier models like when you used to ask them something innocuous about the seahorse emoji.
What I'm arguing is that yes, it's your fault if you misprompt the model and it does something catastrophic. It's also your fault if you prompt it correctly and it does something catastrophic anyway due to some other glitch beyond your control.
The whole discussion started with the point that OpenAI could be financially responsible for damages if their models cause problems for users. If my job is to design a system that works, and instead my system doesn't work, it's my fault regardless of whether I used no LLM, a local LLM, or OpenAI's LLM. Depending on whoever's in charge of doling out consequences, I might get away from it with zero, light, or heavy consequences. But at all levels, I can't reasonably expect to deflect blame onto the model itself.
In all of these cases, it's you. It would be the same if you downloaded an open model, ran it locally, and it happened to make the same catastrophic mistakes. The consequence to the provider is that if they offer a product that does these things, people don't buy the product.
In general, the person whose job it is to provide the company with a working, non-adult website and not hack into other websites is the one who would receive consequences for failing to meet those expectations.
The problem I see here is that ultimately, you'll have capital wanting to replace workers like others have said, and have someone roughly equivalent to a manager or vice president driving teams of agents to achieve business outcomes.
These tools can push out more results than a human can hope to evaluate in a business-sensitive, or even realistic, amount of time. You have to take it at its word that it did things right, and there's no real fear of failure or consequence on the behalf of the agent.
If future jobs are simply reduced to liability scape goats (or more appropriately reverse centaurs) for management to pin things on then I'm taking up goose farming.
That's more-or-less what you are now, especially if you work at a company like Meta where 1) the guy at the top holds majority control of the company's shares and 2) keeps making massive, expensive mistakes either by accident or design.
ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them.
In practical terms? They usually get solved with a bigger badder LLM. "New ideas are needed?" Nah - ten times the params, ten times the test time compute.
ARC-AGI-3 was more of a failure in that regard than -1 or -2, because even on day 0, an off the shelf LLM with a harness could get 50%+. And messing with evals by forbidding "LLM with a harness" from scoring? Yeah no, that was just bad.
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".
"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
We are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic.
Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation.
I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.
If something at rest is accelerating at 9.8 m/s^2, how long in seconds will it take to reach 10% of c? Answer to the nearest order of magnitude - will it take approximately 1000, 10k, 100k, 1000k seconds?
I’m sure you know this is an exponential growth question but have no intuition of the answer.
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).
Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):
- write a well-received book, write a best-seller
- come up with a new company idea, Run that company
- actually have a decent conversation, maybe someday talk somebody out of suicide effectively
- come up with its own ideas or theories that nobody else has presented
- understand the stock market well enough to trade better than an index fund
- be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)
- come up with a theory of what makes games fun, make a popular game
- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
- be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries
- exhibit metacognition (thinking about its own thinking) and self-optimization
- wonder about things
- observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things
To me all this makes the label of AGI completely meaningless.
What AGI has always meant (eg. in 2019) is Artifical General Intelligence.
Artificial -- something made by humans instead of occurring naturally
General -- not confined by specialization or careful limitation
Intelligence -- the capacity to learn, reason, solve problems, think abstractly, and adapt to new situations
Basically, the metric was that any healthy adult human on the planet represents a general intelligence. This has certainly long been reached.
Also some of the stuff you're listing has long been solved as well, such as listing what it knows and what it doesn't know, and what information it would need. Other is just poorly defined: "be able to argue persuasively". AI can certainly write an argument on almost any topic that would pass any University homework in 2019.
However most humans can do at least some of the things given they spend the required effort.
Some problems presented needs a very large context and some are not much solvable (e.g. trading) since market responds to traders' actions, as well, making it effectively an oracle problem (of computation).
On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them. However, brains in nature never stops. Wonder, daydream, sleep, self-evolve, clean up and eliminate memories and views and much more.
But this is assuming the model is the entire story. The original comment you were replying to pointed out that the harness is just as important.
The hardware of human intelligence is not a singular thing that is uniform throughout. You cannot take the prefrontal cortex white matter out of someone's head and say you are holding a person. Much of the parts of our brains that enable much of our intelligence, is made of different specialized stuff. The visual cortex and sensorimotor regions aren't only there for input and output, they are used by the more thinky parts of the brain to do visualization and spatial reasoning. The cerebellum contains billions of neurons making little oscillator circuits and PID-like self-regulation machines that help make muscles do what they're supposed to, but also provide attention and time perception.
Heck, our brains contain language models, that train themselves up based on a glut of data over a span of about 10 years, and then they become more or less set in stone for the rest of our lives. Of course we can learn languages, but the "Critical Period" is a very real thing that produces a permanent architecture for some grammatical structures, or things like the ability to partition a lexicon by gender for faster lexical access which cannot be learned as an adult if your native language did not have gender.
I'm not trying to make a direct analogy, the point is that the language model doesn't need to be fully "generally intelligent" all on its own for there to exist a general intelligence, because the language model can be part of a generally intelligent system, which can do things like form, recall, and manage memories which are by now a standard feature in basically every chatbot.
> On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them
The parent commenter noted:
"if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI"
Harnesses absolutely can enable models to continue thinking about things. And LLMs do wonder and explore weird ideas like daydreams when you allow them to do this.
I'm guessing their defn of AGI is something like the sum total of all humans' abilities? Still though some of those tasks (e.g. beat an index fund) may very well be impossible, and worse yet a lot of those tasks are not coherently defined.
- Amazon is full of AI books, and they're clearly making money. AI has won multiple literary and artist awards.
- Okay, it's not a "new company" idea, but VendingBench is all about ability to run a company
- Plenty of people disagree with you on conversational quality; see "AI Boyfriends" etc.. (and it's not hard to find people who consider it uniquely valuable for discussing mental health)
- "come up with its own ideas or theories that nobody else has presented" C'mon, seriously? Solving a half-dozen hard open math problems wasn't enough there? What the heck counts as "it's own ideas or theories" at this point?
- plenty of evidence that custom models are starting to do well on the stock market, although I'll admit we're a year or so from any solid proof, since you need a track record to really make the claim
- LLMs have been capable of being a GM for a TTRPG for over a year (although like humans, they make mistakes)
- Okay, conceded, but humans tend to take years and large teams to make a game. Even if the capability existed today, it would take a while to actually build, test, market, etc.. - all made much more complicated by gamers being largely opposed to AI art styles, etc..
- "be able to sort through research and come to conclusions on complex geopolitical/sociological topics" - uh... did you mean to say something else, because "come to conclusions" is... like, LLM 101?
- Hahaha, have you met humans? We definitely cannot do that.
- Uh... thinking about it's own thinking is trivial. Most LLMs these days are built using LLMs, so uh, self-optimization seems nailed, too? We just don't let them do it unsupervised.
- LLMs fucking love to wonder about things
- "observe contradictions and ironies in the social-consciousness" really seriously have you actually used an LLM recently? I think you would find it remarkably enlightening.
I don't think that's strictly true, as I can give it a new gui or tui program it wasn't trained on and it will learn it. Unless you're talking about general abilities like sight, but the same is somewhat true of humans.
If you consider the data on which an LLM was trained on to be points on a very highly multidimensional object, the claim is that the LLM can interpolate a convex hull spanned by those points, therefore recovering a subset of consequences attainable from those points. Obviously this hull includes completely novel points that were not present in the initial data set, so the output of the LLM goes beyond its initial training. And yet, there are clearly points outside a convex hull spanned by any finite number of points, such that we can imagine not all possible outputs are attainable using this method.
The claim is furthermore that truly original thinking, the infamous leaps in understanding and creativity, happen by attaining points outside such a convex hull.
It's hard to rigorously verify or disprove this claim. Hopefully this helps build an intuition of why the claim is not as shallow and obviously wrong as it may seem initially.
This makes me realize there is a higher bar we need to achieve with AI still. The ability for the model to evolve through interactions more on a hourly or daily basis. The models are accelerating but inference doesn’t modify the model.
>be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
i'm not sure what makes you think AI cannot do this already. in my experience, this sort of deep research is something AI is quite good at.
I understand that what it came up with sounds impressive (especially since I know 0 about Myanmar), but on the topics I do know about its analysis routinely have very fundamental problems (even this Myanmar analysis has % that add up to > 100). There's a chance it's just parroting the majority opinion on Myanmar, or making stuff up (and perhaps you could ask it to write a strongly worded opinion in the other direction that would sound equally plausible).
For example I asked it to do a full analysis on the AI bubble, and a full analysis on the risks of Glyphosate, and it came up with a lot of things that sounded credible, but within a few minutes of questing was admitting it hadn't even really checked for internal consistency in its positions, and even doing a 180. It certainly was much faster at gathering sources and reading but it fundamentally doesn't seem very effective at creating a consistent worldview.
And of course the funny thing is it says it did a 180 on one of these topics, great, except whatever it concluded will be discarded because it cannot learn. It's just bonkers to me pretend this is AGI, it probably couldn't even hold its own in this very discussion.
Your examples are things that most humans cannot do, or things that AI can already do. For example most humans, even most intelligent humans, could not write a well-received book, run a successful company, or make a popular game. On the other hand, AI can absolutely sort through research, draw conclusions on complex topics, and argue them persuasively. Likewise, I don't know what you mean by a "decent" conversation, but millions of people converse with chatbots daily, so I don't know why you say AI fails to meet that bar.
I fail to see how what you describe is any better than the old autocomplete-on-steroids comparison. Could the human mind learn to spell every word properly? Yes. Do most (or any) do it? No. Does that mean a human can't?
I think what OP was drawing a comparison to is that AI right now could not come up with an award winning novel from the spark of some creative notion and working up from there, as opposed to just mashing together what has already been done and calling it a day.
But you have to acknowledge how uneven the playing field is. The AI has read every book that's ever been written, and can spend hours of compute time in a few seconds.
I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund.
What I'm pointing out here is that these models appear to be intelligent when they really are simply unimagineably knowledgeable. When you drop the time-constraints it starts to become more and more apparent that human intelligence scales better with time than AI does (much in the same way AI can burp out tons of code but make your codebase entirely illegible within a matter of months).
Perhaps to simplify: my notion of intelligence is how much can you deduce with a constant set of starting context
> I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund
I’m not so sure of that - to get average outcomes in these fields it’s a matter of time, to get above average or extraordinary, you need talent/intelligence/taste.
I would be willing to bet that any human for which we spend $100billion - $3 trillion (depending if you want to count single corporations or global totals) on in an attempt to make them as capable as possible would be able to reach all of those levels.
> A quick search suggests that the most expensive education in the world is something like $100k.
I have a kid in an American university right now, and a quick search of my bank account statements confirms that there are far more expensive educations in the world.
But that list is extremely ambitious. Write a best seller, make a popular game, come up with a truly novel theory, consistently outtrade index funds.
That's top 0.001% human stuff, I don't think you can take just any person and get there through education alone, it takes extreme talent and dedication. There's also diminishing returns when spending on education, it doesn't just improve linearly.
The marginal return on education spending decreases fairly quickly, but obviously becomes zero at the point by which there are not enough hours in the day/year/decade to cover every single topic that humans know about - no matter the talent or resources available to the student.
except you are missing the one versus many argument here. sure we could make one human much smarter, could we make endless copies with the same intelligence? no
For $100 billion we could pay ivy-league level tuition for a million people. You don’t think investing that much in education would yield some good research or companies?
Are you imagining artificial augmentation somehow? Purely through tutors or training programs we seem pretty limited. Otherwise billionaires (or even multimillionaires) could have far more consistently successful kids.
Don't the children of the wealthy famously have a tendency to be successful? Or have I badly misinterpreted the last several thousand years of human history.
To really drill down into that I would think you would need to figure out how many millionair children get tutored vs how many get spoiled.
You'd get rapidly diminishing to zero returns after the cost of university a few times over. Every dollar past that would produce no performance gain beyond that.
> You want a computer program to be able to take a single phrase and execute decade long journies?
In my opinion that is exactly the point missing from AGI: the fact that you still need to prompt it. As long as you have to ask for something, is not general.
Autonomy is not the same as general intelligence. We already have all kinds of fully autonomous technologies that are nowhere close to generally intelligent. Plus, at a certain level of abstraction, human beings also need to be “prompted” to some extent by stimuli. And this is the funny thing about general intelligence as a concept: most of the definitions that come close to internal coherence rely on references to human intelligence, a concept we feel like we understand because we all live it all the time, but whose actual nature and structure is extremely slippery.
I don't get it, human employees frequently need to ask for directions too?
They often act on their own, too, and get things wrong a lot. The reason it works is because of all the systems of laws and institutions we have built around humans, not so much because human minds are special.
It sounds like what you're saying is that AGI should have some sort of free will. I'm not sure why you would add that as a requirement. Could you expand?
I think they merely want something with a functioning long term memory. Something that can exhibit growth past the first 5 to 10 human-equivalent hours working on something.
Current LLMs are worse than most dementia cases, reaching "peak domain skill" pretty much immediately.
>You want a computer program to be able to take a single phrase and execute decade long journies?
> Who will be responsible for the outputs and side effects of such a closed loop system?
Itself. That's the point. We can do it. Until it can met that bar, it ain't AGI. That's always been the bar.
"Being a person in all of its aspects" isn't the same as "generally intelligent". The latter is at best subset of the former, and it's also easy to imagine a system that is more generally intelligent than humans, without being a person. See also discussions of the personhood of various animals who are less intelligent than average humans.
I would bet that llms have talked plenty of people both into and out of suicide at this point.
That nitpick aside, I think that's an excellent list. Especially being able to articulate what it does and doesn't know, or how confident it is. That's something that naively sounds pretty simple, but clearly isn't. And it's something humans aren't great at either (see: Dunning-Kruger), but so far LLMs don't even really have the capability to attempt it.
So the goalposts have moved to include continual learning.
In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.
I hardly think it’s fair to label an objection so old that Turing included it (and discussed it at length) in the list of objections to thinking machines in 1950 “moving the goalposts.”
> These arguments take the form, “I grant you that you can make machines do all the things you have mentioned but you will never be able to make one to do X”. Numerous features X are suggested in this connexion. I offer a selection:
> Be kind, resourceful, beautiful, friendly (p. 448), have initiative, have a sense of humour, tell right from wrong, make mistakes (p. 448), fall in love, enjoy strawberries and cream (p. 448), make some one fall in love with it, learn from experience (pp. 456 f.), use words properly, be the subject of its own thought (p. 449), have as much diversity of behaviour as a man, do something really new (p. 450). (Some of these disabilities are given special consideration as indicated by the page numbers.)
Just because a condition is new to you doesn’t imply moving the goalpost. People have been putting forward continual learning and similar conditions like autonomy since 1950s.
Ordinary people do these things all the time. There are new companies made every day, new books top the charts every week/month/year, same for music.
People have decent conversations every day. Ordinary people sometimes do have to talk someone out of suicide.
Yes, average humans are not beating the stock market. But the average human is a bit better than you give credit to.
Please consider the context of the question. An artificial intelligence only needs to have the cognitive abilities of a random average human in order to be "AGI".
The average human has never published a bestselling book. A person who has published a bestselling book is an above-average writer. And, therefore, an artificial intelligence capable of writing a bestselling book would be above an average human at the task of writing books. Therefore, somewhere beyond an AGI.
Attempting to redefine AGI to "being better than most humans at most tasks" is moving the goalposts towards artificial superintelligence.
It costs money to train each single human, who is then only productive for a number of years until age takes its toll. Once you have trained one software system, the marginal cost of producing a copy approaches zero. Every subsequent improvement can be broadcasted in a matter of seconds across thousands of data centers. In addition, software does not get sick, age, or die.
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month?
I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and the ability to prioritize correctly, i.e. tell which of the n things it could do next is the one that actually matters.
Can't say for sure that those are enough, but not having them seems to be most of why I still have to "babysit" these incredible tools.
I agree, we're not at that kind of long horizon capability yet. You still need a human in the loop to do manual testing. For whatever reason, we managed to automate the skill before we managed to automate the focus.
So, for now, humans need to stay in the loop and do low skilled labor to keep the skilled work the models do from going off the rails.
Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.
Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.
And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.
At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.
But sure, they can create a decent website or CRUD app, so they must be really smart.
But that happens with humans as well. You are having the same experience with an AI that many managers have with their direct reports.
The smarter AI gets, the easier it becomes to move the AGI goalposts. Seems at this point there are people who will refuse to call anything less than omniintelligence AGI.
(And then the excuse will be, but it’s not omniscient! And even if it were, is it omnipotent?)
My success rate for solving software engineering challenges encountered in my day jobs has been near 100% for my entire career. I can only think of a few tasks I kicked back and said they were impossible. For example, after trying to get a signal processing system working reliably I decided to sit down and calculate the actual limits of the channel we were sending the data over and found that from a basic estimation it would not be possible to do. In start ups you don't really get to get stuck in a spiral and not fix things.
I find agents often get into these cases during research tasks.
Yep, people are typing comments with a computer that is powered by several layers of software that will be stored on another computer powered by several layers of software to be read on a computer also powered by layer of software. And then they hope to make the argument that humans cannot produce software.
yeah for me it's "I wanna add this new thing to an existing system" and the AI responds "we should just add some arbitrary state here to facilitate this feature".
The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this, The AI however knows the shitty solution would solve the immediate problem because it's been trained on shitty solutions.
the problem could simply be the AI doesn't have all nebulous loose context I have about the goals of the project and future plans, but I would have to write a novel to give it that context.
> The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this
1. I’ll often include boilerplate in a prompt to tell it to make the broader fix. [1]
2. However, a top HN AGENTS.md post 11 days ago included the standard guidance “As much as possible try to minimize the number of changed lines when implementing a feature.” I.e. some devs want LLMs to avoid broader changes and so some of that likely makes it into the training, even if others like us want the opposite.
[1] As far as whether my boilerplate is effective, I don’t know.
Humans dig ourselves into holes as well. Sometimes more intelligent humans are better at realising they are digging a hole and clamber out, but sometimes they just dig deeper.
And: Is your work more difficult than finding proofs of or counterexamples to decades-old open problems in mathematics?
Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.
Current AI is arguably much more capable of multimodal output than humans. It can produce an incredibly vast variety of audio, images, and video. Humans are limited to producing the sounds we can make with meatflaps in our throats, and contorting various parts of our bodies to produce crude symbols and shapes.
(Very capable!) Embodiment, persistent operation and continuous learning are indeed things that still set us apart from AI. None of those are fundamentally difficult to solve, though.
More importantly, none of those are particularly relevant for being "intelligent": If a criminal threatened to kill your family unless you solve some difficult problem that requires only intelligence and you could choose any single person, animal, or AI to help you with it, which would you choose? Be honest.
That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.
What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.
I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.
Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.
If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly.
(Edit: I wrote ARC-GIS the first time around, for some silly reason)
It's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.
Surely part of the problem is that intelligence seems to be implicitly conditioned on embodiment, to the point that "covering only cognitive tasks" seems inherently ill-defined or arbitrary. Everything we do is a cognitive task. At some point, the criticism will be "sure, it can solve research-grade math problems, but it can't fold my laundry".
Even our large language models have an implicit embodiment in the domain of text (and more recently, multimodal inputs). That seems sufficient for certain things, and insufficient for others. I suspect that AGI that does everything a human can do eventually turns out to be fairly analogous to humans in terms of sensory input and domain output, even if the scale is radically different (e.g. thousands of robots uploading (touch, sight, audio, smell, etc.) sensory data to a single model, and each being actuated individually).
>AGI has a pretty precise definition, covering only cognitive tasks.
OpenAI's own charter defines AGI as "Highly autonomous systems that outperform humans at most economically valuable work". This is actually fairly sensible and involves obviously a ton of non-cognitive, emotional, social and physical activity. In other words, if you can replace most or all human beings with a machine, you have something that's generally intelligent.
That's obviously not even remotely where we're at, AI chatbots do well on narrow usually text based or programmatic problems, but can't even replace a barista or a plumber.
There is rapid progress in generality in humanoid robotics though. I think within the next year or less we will get the ChatGPT moment for humanoid robots. If you look at progression of capabilities such as the recent Skild AI demos.
Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
I think I have the following questions about what AGI would look like:
1. Do you expect an AGI to be able to competently do any knowledge work that an able human does today?
I think this is implied by the "General" component. I would assume that anything we call AGI would be able to do any of these tasks if given the time and reference material needed.
2. How would you expect AGI to handle edge cases? (Missing context, no known solution, under specified instructions, over specified instructions)
I would expect an agent to be able to look at the context that work exists in and correctly attenuate it's intentions for these goals. Simpler solutions, more thorough reporting, etc based on the need.
3. In my work I attend meetings, write reports, write code, research things, etc. Would AGI be able to reliably do that?
I would say that AGI would need to do this. I would classify this as the "Intelligence" component. Obtaining context, building a model of a problem, solving it, and convincing others.
4. Would it be able to inspire trust in itself? Trust can be established through verification of it's outputs, the construction of introspective tools, no hallucinations, etc.
I would say yes to this as well. It would be a component of the "Intelligence" to know that buy in is more important than the completion of a task.
To these points, will Astra be able to do these things? If not, I would hesitate to call it AGI.
To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.
Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
Whether AI is AGI does not depend on the speed at which it operates/thinks. Clearly all the theoretical work done by AGI will be done orders of magnitude quicker than humans can do it.
It is an open question to what extent practical experimentation/work will be a bottleneck for the theoretical work. It stands to reason that it is improbable that it will be the bottleneck for 100% of the speed of treatment development.
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
Plagiarizing on a massive scale to generate works which appear to be Fields-medal-level results is not the same thing as inventing new conceptualizations in mathematics.
No matter how bodly they write the headlines, what has happened in mathematics using Large Language Models is very much "inference over a bounded search space" even if those bounds are immense.
For a comparison of true creation of novel conceptualization in mathematics is submit the works of Martin Hairer, one of which is Introduction to Regularity Structures, [https://arxiv.org/pdf/1401.3014] None of the so called, "novel math discoveries" by any LLM is as enlightening and expands the state of the art in math like any of his writings.
Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).
If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.
>what would make you think Astra is yet to be AGI...
Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI...
And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are targeting....HN appear to have fallen for it, as usual...
They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.
The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.
The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.
> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.
You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability.
I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.
> You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes.
Yes, and I could script a truly marvelous proof if this textarea were but a little larger :)
Hand waving doesn't count for much these days when you could spin these things up quite quickly to prove the point, so the GP's claim seems much stronger than whatever you're not convinced of?
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Simple. AGI is undefinable and benchmarks are notoriously flawed.
> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.
People are ultimately capable of self deception. Believing that one is 'generally ... substantively above average on human benchmarks' may be more indicative of the brittleness of the claimant's human benchmarks.
Your observations expose the brittleness of the benchmarks being used for Fable, where the 'reasonable confident' claimant is working in the problem space of those two benchmarks, and you're demonstrating the failure at depth of the very capability facade for the model.
Sure, both the model and the confident human get the first layer right in some benchmark, but at least the model, and probably the human as well reach a collapsing probability of accuracy quite quickly. They may be completely and unconditionally right that Fable is more intelligent than they are, but that has little relevance at intelligence in depth as measured against your hobbies and the risk of unsafe or false responses.
Being smarter than a select group of people under test conditions is not conclusive regarding AGI. Moving the goalposts to declare AGI via a shallow benchmark, when a model could be proven wrong by almost anyone with basic competency given a few turns of iterative depth is where the 'grift' resides.
We're so far from anything approaching actual AGI, and it is highly speculative to infer that the current approach to Machine Learning as applied in LLM is even on a path that leads to AGI. But sales and promotions teams gotta pose, and we can all hate both the game and the players.
FWIW I believe we can hit AGI! but I think at this point it’s clear that benchmarks are ~meaningless. LLMs are spiky / alien intelligences which don’t map to our own expectations; the existence of a benchmark creates a dataset to hill climb & RL is really not generalizing well.
I’d go out on a limb and say astra’s ability at graduate level math will have ~0 bearing on its general reasoning capabilities; we’ll all acclimate being tired of its “neuralese” and more surprising mistakes.
I think we need a true, step change advance in model architecture, but it’s hard to see how the current frontier labs can do that because of golden handcuffs / innovators dilemma
AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
I agree that LLMs are unlikely to be the final form for AGI, but what you are talking about is orthogonal to the IQ and for most cases general utility. It's like looking at a savant chained to a workstation reading tasks from a conveyor belt and saying that it will never have human level capabilities.
The model doesn't "know" how to generate tokens any more than it knows how to stop generating tokens. The sampler simply stops pulling values when it outputs a "stop token", which is a token the same way every other token is.
That is to say, it stops when it's statistically the most likely to.
An AGI test should be black-box; we shouldn't impose require requirements on internal components. As long as the overall AI is capable of learning and remembering things, it shouldn't matter if there's a stateless LLM internally.
I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have.
The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints.
But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad.
All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks.
Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
To me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake.
A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following properties: one shot learning - fully (and always) online, perfectly efficient (through self improvement), no context limitations ie. persistently thinking, not just awaiting input.
So for me, no, not AGI yet. But still very intelligent and capable (and perhaps it’s safer this way?)
Sentience and intelligence are different things. Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!).
Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).
A model that can't beat gemini flash 3.8 on deepSWE is not AGI.
I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
I feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges.
I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.
The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).
If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.
I'm still not convinced we've passed the Turing Test.
Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?
I agree that it's misleading but harness is now an essential part of LLM's effectiveness. It's safe to assume that LLM-alone-AGI is not coming anytime soon, given most of the frontier LLM vendors are developing their own harness.
Also the training dataset is proprietary and they'll drive the LLM's behavior, so it make sense for the vendors to invest in the harness and bake in prompts that work best with their models.
In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.
Per the Fireship video, it was less that the answers were in the training set and more that the ability to calculate the flag on Exploitbench was left in from the previous test that went awry.
I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.
Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.
I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.
So I don’t know why it can track fib algo, but no chess concepts.
- an undoubtedly very intelligent person. in the course of their studies, they have read about different chess strategies, openings, etc. but they never actually played the game themselves
- an average person with a year of chess playing experience
who do you think is going to win? of course, you could give the LLM time to think and consider its opponents potential next moves, but this is a computationally expensive way to play the game that doesn't scale
which is all beside the point that chess isn't a very good proxy for general intelligence. there is a correlation, but it's very weak
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Can Astra, or any other model explain how exactly it reached this or that output result? Start with a simple query of asking to add 55+66 for example. (no LLM program can do that)
Can Astra, or any other model refuse to answer or go on "thinking" in a orthogonal direction on it's own?
That's just two quick ideas, I'm pretty sure cognition scientists can invent better and wider range of checks.
In my opinion this is goal post moving. Humans do many things that we cannot fully explain either without post decision rationalization, and not all intelligent humans are deeply introspective.
> The ARC-AGI-3 scorecard is extremely misleading (...)
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
> I am reasonably confident that there's essentially nothing that I am better than Fable at
Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.
And my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous.
It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?
Are you using this satire to argue that a benchmark self-labelled AGI is as scientifically rigorous as climate change data, and not just a random marketing decision?
If you think they're announcing AGI as a marketing decision, you are blinded by the accidents of your birth. Capitalism is strong -- humanity's instinct for communal preservation is stronger, sometimes.
And yes, the one deeply-researched field going back 75 years is as scientifically rigorous as another deeply-researched field going back ~100 years. I guess you can draw climate studies back to Descartes and the Islamic golden age, but that doesn't privilege it in a time where the methods have changed completely in the span of decades.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this.
Makes no sense.
We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism.
I am pretty sure the average human would not have done this (among other slightly less absurd examples in the article requiring employees to fix it's mistakes):
Andon Labs added: “During the first week of operations, Mona purchased 120 eggs despite the café having no stove and to solve spoilage issues ordered nearly 50lbs of canned tomatoes intended for fresh sandwiches. Employees eventually created a shelf displaying Mona’s strangest purchases: 6,000 napkins, 3,000 nitrile gloves, industrial trash bags and 2.5 gallons of coconut milk.
Sign a contract? Learn things over time and retain them?
Mind you, the original thoughts on AGI before Sam Altman started to water them down involved continuous learning, which LLMs do not do, their core data is static.
Well signing a contract is more about bearing responsibility, even if you granted LLMs "personhood" they cant' meaningfully bear responsibility. So unless OpenAI is ok with having their C-suite face every consequence for what their agents do, including jail time, fines etc, then it doesn't matter.
I don't trust the people who are building the models no. Ideally, if humans were not so broken and untrustworthy, then yes. I want the Jetsons future of robot maids.
AGI does not ever have to be achieved. It is enough that we (as a species) persue it, and continue moving the goalposts each time we learn something new about the limits of our technology and how to express those limits. Because that will progress the technology, no matter what we label it.
I am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Stick to the original definition of AGI of an AI model being able to self-improve independently with 0 human intervention and become an "everything" solver. Ever since money got involved in this, the goal posts have shifted considerably. If OpenAI truly had an AGI on their hands they would then be able to crack encryption, destroy world markets, and funnel all resources back into their new for-profit organization. Since their mission is now share price, until I see any evidence of an infinitely growing stock I will reserve my congratulations.
Using a harness designed for a specific problem set to solve that specific problem set, means the AI+harness is generally intelligent? How do you figure that?
Or do you mean that, for any given problem, we could theoretically design a harness that allows AI to solve it (not that, one single harness solves everything). In which case I'm still not convinced but I guess could see why one would believe that.
Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
arc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today.
Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
AGI is a meaningless term that can mean nothing and everything at the same time. It can be used by AI bros to hype their latest releases which are always one step away from achieving AGI, or it can be used by anti-AI people to say it's not AGI because of X arbitrary thing they decided on in the moment. It's a term of pure convenience meant to obfuscate other more pressing discussions on the topic.
Most telling is M$ or whichever one of these borg megacorpos defined AGI as (paraphrased) "AGI is whatever tooling earns us a gazillion dollars in revenue"
With how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
I might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
My whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
Turing never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.
Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.
That's a very interesting thought that I hadn't had before: what would Turing think of where we've arrived with machine intelligence? What would be his approach for testing?
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines.
I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
There's a lot of truth to this: Execs want their own use case covered. In addition, everyone wants to be better than the competition at basic use cases, because that's what people compare first - even though we all know they are not a good reflection of reality.
I'm currently working on exactly a project like this.
I feel like there are so many cloistered people at these companies that they are left scratching their heads about what normies even want. Like, they literally can't fathom basic stuff that isn't just highly consumer-oriented. I dunno, like applying for government services, paying your gas/elec bill without being confused af, keeping the dr up to date with your dad's illness, or how to get your newborn to sleep at 2am.
Many VCs also dislike these examples, I believe. I'm doubtful this is what they're being pitched.
As for public releases: I wonder if it's because these examples are easy to relate to. Many websites are just a long tail of industry or use-case specific stuff. What's valuable to me probably means nothing to you. This is unlikely to resonate with people-wit-large (and LLMs are marketed broadly) or requires the reader to think (and marketing that requires thinking is bad these days).
Second, it's arguably a good litmus test. If it still can't do the worn out examples of plane tickets and shopping, which would be a good assumption since we've been demo'd these use-cases for 2 years at this point, then ...
Maybe because the people making them are workaholic types who really don't care? I've certainly been in situations where I didn't really care what shows up for a meal. Someone was tasked with getting food and we leave it up to them. Sure, I can imagine such a meal being bad and it has been a few times in my life but 24 of 25 times, maybe more, it's fine. Further, the AI knows your preferences.
That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
When I tell a bot to find the best value per volume for a reasonable quantity of unscented Dawn dish soap [so I can buy that], then: It often makes a complete mess of this seemingly-simple operation.
(And yeah, that is an actual thing that I've tried to accomplish with voice commands while standing in my kitchen and doing some dishes. It seems very simple, and it did not go well.
Maybe when we get the basics figured out we can start worrying about how inept it is at doing vacation planning.
It seems that this kind of thing isn't sorted at all, and that this is a very real problem for those who are in the bot business: These missed opportunities leave money on the table.)
Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.
And I could bet money on that in short time after someone builds that sort of system the next step is to make it worse. Push worse and more expensive options to user. Or at least those from highest bidder... Anyone involved just can't keep themselves honest so it is doomed to be exploitative.
I think it depends on what you do for work. I'm not going to ask an agent to book my flight for my vacation to French Polynesia. I want to pick my seat and potentially find a deal making an upgrade worth it, choose an airline, etc.
But my routine business trips in the CONUS with strictly defined booking options... let me just email an agent "Get there by meeting on day A, leave after meeting day B" and have it sort it all out without the drudgery of the corporate travel portal. YES PLEASE!
By that point, you don't need an AI that boils the oceans and hopefully doesn't confabulate or misinterpret your words, all you need is a better corporate travel portal, which is a legitimate workplace productivity discussion to have, and a solved problem with traditional methods. Ours is effectively close to the workflow you describe: picking dates and time brackets, destination, fine tuning flights and hotel options, sending for validation, less than 10 clicks through information-dense and effective/predictable screens which I wouldn't want to trade for a chatbot and it's usually over the top words salad.
We have had idea of expert systems for decades at this point. And spend however much on software development. Some how we have failed to reach this in way too many places. Even simplest things like cancelling some service might be broken...
And now we want to add some random factor into middle of it all...
Plan a holiday, definitely. There are many esoteric things one has to research to properly plan a holiday that I’d rather just not. Some things require reservations months in advance and I’d rather an AI just figure all that out for me ahead of time.
For me, it is not a matter of trust but that I actually like shopping, planning a trip, deciding what restaurant to go to. Deciding what to buy when shopping is a matter of personal taste and not intelligence.
A human assistant is largely a status symbol. Most people are not really that busy. The real problem with an agentic assistant is if everyone can have one then it no longer acts as a status symbol.
"Somebody is wrong on the internet" will get you more, and more reliable, results on what people actually want from AI. You are now participating in a very high-value survey.
You can't even get many people to buy things online at all and if you can it's less profitable than retail, because you need to spend a lot of money to convince people, advertise to be seen, and account for returns. I think this is also due to the factors you mention.
One quick example: In fashion, Inditex and Shein have about the same revenue (€39.9bn and $41.8bn in 2025), but Inditex is more than three times as profitable. I don't see how there is a demand for agentic commerce that would remove even more control from the customer when shopping. Part of why we shop is for the experience. For B2B producurement platforms like Alibaba I can see the appeal though.
I recently needed to buy some hardware for a piece of furniture.
Ran Codex, it found it for 18% less than what I found in the top Google results. It did it by finding smaller shops, applying a discount code, subscribing to a newsletter for a better code after approval, and took into account the shipping (by placing it in the cart and going to checkout) all to get me the best price.
I’m guessing without it I would have spent much more time on it and paid the original price I saw.
If you use AI agents well, they can easily save you more money than they cost, and saving money is something most people are pretty excited about.
I mean, that's cool and all, but the numbers are really going to shift when it accidentally goes off and orders that same hardware from every vendor in your local region and the top 5 online results for comparison.
It's the same problem as all other LLM solutions (that I hope OpenAI is working on!) it's non-deterministic, and there's no way for the user (or model provider) to know what the distribution of possible outcomes is. This just gets compounded when multi-call harnesses come onto play.
I‘m using mostly the browser automations for things like that now. Same for research - ChatGPT is banned from reading many pages, but Codex can read anything I can.
But isn’t it funny that Cloudflare is blocking AI on their pages, but on the other hand is researching and marketing things like „you can put a browser in a CF worker“
Their browser workers won’t get blocked. Same with all big vendors, keep out small competitors enjoy access yourself and sell it to a select few partners.
I expect a small site that undercuts the top Google results by 18% with a sign-up discount probably isn't profitable on those orders, so blocking agents would save them money - it's not like someone using agents in that way is going to have any loyalty to shopping from that site in the future.
And yours and other's agents will probably remain an insignificant and invisible customer base anyway.
My crystal ball is as good as anyone's, but if "agentic shopping" ever becomes mainstream, you can be sure that the vast majority will ask their phone (i.e. Google, i.e. Google Shopping) what the best price is anyways.
Sure, Google could be their agent, or ChatGPT, or Claude, or whatever local model the person is running. The big guys might all have the results cached so they don't have to rerun the crawl. Whatever their choice of agent, it seems pretty clear that almost everyone's going to use them, they're way too useful not to.
The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint.
I wouldn't want it to pick food for me from a place I've never been, though to be honest with enough order history it could probably do a decent job at it.
I work for a larger german retail chain and agentic shopping is already on the "near future vision".
No one thinks this will be used but somehow shareholders love it.
Agreed. There's not many things I don't want AI to help with, but buying stuff autonomously is high up on the list of things I don't want. Brockman's latest interview was something like: "AGI would be able to say oh this band is playing, I bought the tickets for you and arranged your flights - I hope you don't mind" (paraphrasing here). I definitely don't want AGI running my life like that so I can be a mindless consumer. I'm sure the advertising/marketing companies would love it though, so they can make closed-room deals with AI providers to shill you garbage you don't need. Just another reason why open-weight models need to keep up.
True, but they're still friction to be reduced here.
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate.
"Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip!
I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming up with such demos.
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
I'm still here, nearly 50 years and counting. If you had asked me what I imagined AGI would look like back in the 90's, I would have told you "A system that can do everything we can: see, hear, think, do.". If you had shown me GPT-6 back then, I would have said "It looks like a really powerful program, but that's not really what I had in mind.". That's AI, but it's not quite general.
As someone who spent countless nights tweaking Edge Detectors (looking at you, Canny), morphology operators, etc., building models to recognize 10 handwritten digits, let me tell you: the current set of LLMs (even the smaller ones) seem like magic. I had never imagined a computer would do such things in my lifetime.
Exactly people can say whatever they want, but current level of LLM is AGI level to me. It is already on par with senior programmer if the instruction/prompt is right.
Once we have 1000 tps, i am sure robots etc.. will also start working like magic.
I don’t know, I am writing a modest 30 page paper with Fable and even after rounds and rounds of feedback and improvements there are so many things that are just plain wrong or weirdly out of place or just stupidly written that Fable 5.1 doesn’t seem to have any awareness of by itself that I don’t think it’s AGI, I think a human researcher can easily outclass it in writing and problem understanding. It definitely has super human capabilities but it lacks awareness or self reflection in my opinion.
For example it should be easy to tell it to not write a paper in the style of a clickbait SEO article or use all of its stupid hallmark AI writing patterns “it’s A, not B!” And a smart human that would be told that would be easily able to comply with that but the model needs to be told in a very detailed way and it seems to lack even basic capabilities to reflect on this, when explicitly given a sentence it will be able to rewrite it but otherwise it’s mostly blind to it. That’s to me a hallmark of it being overtrained on the specific tasks or problems so it appears very smart but once you go off script it still shows that it’s not a “real” mind.
Of course it’s amazing and has super human capabilities in many areas but if you honestly think it’s better than Einstein like some people suggest why can’t it write a simple “good” academic paper even after giving it specific examples and instructions.
Maybe that’s what makes these things dangerous, they have super human capabilities in some areas but apparently lack self awareness, taste and meta reflection abilities. The only reason people aren’t afraid more is that they don’t act in the physical world yet, imagine giving it a body, superhuman strength and letting it care for your child when it has a strong “urge” to comply with your exact request and little to no self awareness and human basic instincts.
I believe so. AIs are shockingly good at a lot of domains, but there's still a lot of pretty basic stuff they don't really "understand" at a conceptual level and (currently) they can't learn to get better at them.
(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)
If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.
True. "AGI" has also become a marketing term. Achieving AGI has become valuable, so companies will move the AGI goalposts, over and over again, so they can achieve AGI, over and over again.
If you came at it from the perspective of imitating what the human brain does, we now have a very very powerful speech center and short term memory, and vision catching up. The other parts are missing. I‘m sure that’s being heavily researched.
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
I can have Astra run a large-scale infrastructure migration 24/7 (much of the time waiting for results), completing it in weeks or even months faster than I could before agentic AI.
For whom? That is a fantastically ill-defined test. Everyone here is comfortable throwing around this or that is or isn't AGI which is fun because, at the same time, nobody seems to have a testable definition.
If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.
We've had AGI (artificial general intelligence) probably since the first release of ChatGPT, and certainly since the first agentic harnesses. They're just finally acknowledging what the term means.
Artificial. General. Intelligence. The ability to solve (even partially or even badly solve) problems drawn from arbitrary problem domains without pretraining on the specific problem class. You can pose any problem of any type using natural language to an LLM and it will attempt a solution. That's literally all the term means.
You (and the rest of the media and many industry figures) are conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).
Reference class in this case means not an example but what the comparison is against. Superhuman means better than humans. General intelligence is defined without any reference to human capability levels.
This is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,
I think we're getting to the point where it is difficult to identify the goal post of AGI.
Is it rapid skill acquisition? -> ARC benchmarks are saturated
Is it breadth of knowledge? -> See many ... many benchmarks
Is it ability to do hard tasks? -> see terminal-bench and released outputs.
We are at the point where the starting point for most tasks should be "send your agent to work on it."
So where do we draw the line in a way that doesn't move every 6 months?
The real answer is converting from any format to any other reliably. Text to speech, speech to text, music to video, image to 3D, piloting a drone by converting video feed to rotor speeds, literally any file conversion, like html to pdf, photoshop project to png, png to photoshop project,... turning Toy Story 1 into a series of Blender scenes with all textures, models, materials, lighting, camera movements matched to a tee, should solely be a matter of how long you let the model run. It should never run itself into a dead end. It should instantly know when it is making mistakes, with no human babysitting it.
I can do none of those things.. I hope that I am generally intelligent.
1 year ago we viewed models as tools and agents were just kinda toying around, that we now think the bar is literally an anything to anything converter through one agent is wild.
It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.
If a video announcement and a press release would change a person's mind on whether this is AGI, I don't put a huge amount of weight on that person's conception of what AGI is.
There has stopped being a formal procedural consequence for OpenAI leaders to declaring AGI, there is a clear (small) business benefit to doing so, and the capabilities of all the frontier models are impressive. So why not declare AGI? It's not like anyone can prove it's not...
Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.
> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
I don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Yes, the very explicit plan of both OpenAI and Anthropic is to use the not particularly efficient LLMs to automate their own AI engineering. That seems to be going well - on coding front and model tuning front so far. They have more planned.
And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.
It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.
I recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code.
But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.
What about finding the 1st known complex structure over S^6, proving Ehrhart’s volume conjecture, proving a sharp "density" bound on primitive sets conjectured by Erdos >60 years ago?
> Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
It's not good to be confidently wrong the way you're being.
I mean, the plan is to use these models to find and solve those gaps. That's kind of the whole pitch of these companies: they spend a TON of money upfront setting up this infrastructure, but each iteration yields a system capable of making the next iteration even better.
>The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Perhaps. But only at that point, not leading up to that point.
It's kind of like setting up scaffolding to build something. You spend all of that time and money to build something just to tear it down in the end. But the point is that it's simply a cost to be able to build the actual thing you're building.
If these companies are able to achieve the results they're looking for, none of the investors involved are going to care that the datacenters and infrastructure they spent so much money.
If we manage to get to AGI and it looks, works and behaves like a human brain... I mean, cool, but that's a very useless AGI compared to the incredible stuff we have access to today.
The HN crowd I'm sure will still be unhappy calling it AGI because "it's not AGI unless its speech comes from the cerebral cortex region of the brain, otherwise it's just sparkling emoji" or something.
I think the idea is that you wouldn’t need humans to do anything anymore, right? As impressive as it is, it’s still ultimately directed by human planning and coordination. Assuming they are aligned, you could have a collection of AGI that you let loose and they tirelessly solve all of humanity’s problems, do all of our work, and progress science and our understanding of the universe.
Those are all things that humanity is doing everyday. What we have is amazing, but it’s not that.
>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
I think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.
(kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice))
Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to continue to increase their market share from Anthropic.
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
Pretty efficiently, apparently, since it saturated ARC-AGI-3 in half of the predicted time, and according to the Chollet blog post on the fly created dense DSLs to describe and analyze individual games.
I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute.
Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
You can argue that TSMC has no moat since Intel and Samsung are also able to eventually make a node as good as TSMC - just a few years later and at smaller scale.
No. In the semiconductor industry, the "catch-up" player isn't normally spending less in absolute R&D terms.
Comparing the R&D costs of creating GPT-4o vs. DeepSeek V3 (the latest gen for which we already have good accurate numbers) it looks like the latter cost 1/20th as much to create.
If Samsung could catch up with TSMC for 1/20th of the cost, people definitely would say that TSMC has no moat.
That's the ratio the widely published numbers give [1]. One does not have to believe the numbers [2], but those who do believe them are then justified to conclude that there's no moat.
Which numbers you believe is of course going to affect whether you think there's a moat or not. That's largely orthogonal to your TSMC/Samsung analogy I responded to. If you think the "moatists" are wrong because they believe the wrong numbers, that's fine, but then there's no need for the analogy.
I am not a "no-moatist" per se but one can argue their might be a plateau to how good a inference llm can become. If this is the case the playing field shifts to context, tools and harness, which are much cheaper to build an compete on.
Ultimately, that's what I need to be convinced. No one has put forth a good argument yet.
Clever architecture --> Ok but OpenAI/Anthropic can use these as well and they also have very smart people with their secret clever architectures
Distilling --> Ok but distilling means you will never be smarter than the original. Furthermore, reasoning is now hidden by private labs and they have poison pill answers for distilling if they can detect it. They will be able to detect distilling better and better.
Cheaper electricity --> Ok this is cancelled out by their chips being much less efficient due to not having ASML EUV machine access.
So I don't see why fundamentally their training costs are cheaper over the long term.
Yeah, I'm not sure if "no moat" analogy stands for chip manufacturing. Even if foundries acquire lithographic nodes, the procedures (temperature, duration, etc) are for them to figure out and are usually kept secret. This secret could be the "moat" that differentiates each foundry's operational capabilities.
About the same, 5-10, when you consider major (aka frontier) airlines.
Actually not a bad comparison. Both burn massive amounts of up front capital to protect an oligopoly in the hopes their commodity product eventually pays off.
I think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).
It's not energy costs. The US produces about 70% more electricity per capita. Chinese households do pay less than half what US households pay for electricity, but that's because the NDRC sets prices below costs for households. They make it up by charging industry more, and the industrial electricity prices in China are roughly 34% higher than in the US.
They have models for that. That's what the Nemotron series is. Not just open weights but open training data too and full tutorials on how to use them to fine tune or train your own models.
They exist to keep people using and advancing the tools on their hardware.
How? Imagine an open-weight model comes out that is somehow better than proprietary solutions. Now the marginal cost for the consumer is just the cost of renting the inference hardware, without having to pay the overhead of the owner of a proprietary model. And because it is cheaper, more customers want to use it, and Nvidia will sell the providers the inference hardware that they need.
1. No open ai and anthropic means no buying gpus to train. Now nvidia spends money on hardware training their own models. Opportunity cost plus expense.
2. Any open models created from this will not necessarily need their silicon, see apple mlx.
1. I don’t think that’s a very strong argument. OpenAI and Anthropic don’t buy the vast majority of GPUs they use they rent capacity.
Nvidia could just the same rent those GPUs out for inference and actually have way better margins than they do right now. Antitrust and putting all your eggs in one basket are why they don’t, similar to TSMC.
2. Neither do AI labs. See Anthropic buying TPUs, deploying with AMD. OpenAI on Maia, Cerebras, their own wafers.
They have a lot of moat, i'm not sure what youa re talking about. Only amatures are using Qwen, open source stuff that is 3-8 weeks behind. Plus OpenAI has some verticals that keep people in there.
This is perpetually an issue with the whole field of AI/LLMs. The experience is so personal. Every time I talk to someone about their use of LLMs for software engineering, I'm shocked by their approaches and experiences. They say "X model keeps missing things" when I rely on it heavily for being thorough. They say "Y always gives me the best results" when I can't stand it.
People will see/think that I'm doing very well with my LLM use, and ask me what I'm doing. I tell them, they try it, then later they come back to me saying they just couldn't get it to work.
It’s really inconsistent. There are sessions where it nails everything perfectly and I leave happy. Then there are sessions where every turn it corrects itself and changes it mind. One session recently I found it funny how every single time it did this one task it tripped over itself and killed its own connection. Like 20 times. It didn’t bother me I just found it odd how despite it being noted down in its state file it kept doing it over and over like some idiot. Literally they can’t learn from their mistakes yet.
I’ve found Sol performance to be incredibly spiky. It has tremendous IQ and can fix very difficult bugs. But it is horrible at design (both visual and system design), anything that involves thinking about users or UX, and massively overcomplicates almost all work.
I vastly prefer Sol. It does what I tell it to almost exactly, pretty much every time.
I work on very low level stuff (think RTL/FPGA, firmware, software where optimising for nanoseconds is just normal).
For me Sol is the only cost effective model available. Fable 5.1 is indeed good and vastly better than original Fable (which refused to work on most of my stuff for 'safety' reasons).
It's very good at this sort of low level stuff to the point that I really can't understand/relate to people having a good time with Opus (which comparatively performs extremely poorly on my particular workload).
I also just don't like how lazy Anthropic models are. They will do 10% of what is asked and then summarily declare victory.
Sol on the other hand is more like "one of us", slight touch of the 'tism, extremely pedantic, will go to the edge of the known universe if that is what it takes to prove/fix/build what you asked for or run out out of credits trying.
It's a personal and workload dependent thing. For me right now Sol for 99% of stuff because Fable 5.1 still burns through $5k in credits a day.
Can confirm this as well, mostly VHDL and HLS. Sol and Fable can reason about performance and designs consistently. Whereas Opus and others seem to just throw generic optimisation techniques at the wall unprovoked (while hallucinating a justification + expected improvement) until the synth reports improve.
Agree 100%. And I also work a lot on lower level / systems stuff (including RTL here and there, too). Opus is sloppy, and leaves negative cases all over. The GPT models in Codex have a more pedantic and detail oriented "personality." Often to a fault.
Sol will leave a mess of excessive redundant tests and isn't so great at abstraction ; but it produces more reliable working systems.
It's kind of nice to have access to both, but I don't have the $$ for that right now, so I just keep the Codex sub
I noticed the same. I wanted a simple crud webapp and suggested an insane techstack involving C#, Razor Pages, MSSQL and more. I went with my planned setup of python flask with an sqlite db which served me well for years.
It's still incredibly important to have a human in the loop correcting design decisions and having good taste.
Was your prompt just "I want a simple crud webapp" and that's the extent of it? There's absolutely no way you included the words "python", "flask", or "sqlite" and it still went with a Microsoft stack.
We're not asking the model to simplify something, we're asking it to perform a task. Its subtle preferences show up as an overcomplicated path to the goal.
In some cases, there are also nuances that we don't pick up on. Here it's our preference for simplification that's showing up. We set the lossy compression factor higher than it does.
Its funny, my experience with Sol has been awful. It really overworks problems and tracks into areas it does not need to...
I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.
It is funny indeed, people sometimes with same amount of experience with software development, get vastly different experiences from different models and harnesses.
> I just dont get how its good for some, and bad for others.
If I were to listen to my hunch, it would tell me that it's all up to the prompts that ends up going over the wire (including all the bloat some people have), what workflow/process you use and what the existing state of the project is.
You have to bake the 'lazy dev'/'keep it simple stupid' mentality into your AGENTS.md and / or the skills you're using to design things. It will take things too literally sometimes so you also have to make sure you're being accurate. Best way I've found to use it is make it ask you clarifying questions about what you're trying to build and have it help design the shape of the thing. Then it writes the instructions in a format it understands.
I've had Claude do the same thing where it goes off and spends 100% of my tokens on 3 functions and an ungodly amount of tests / scaffolding that do almost nothing when I gave it an underdeveloped idea.
Codex is missing a few things that Claude code has had for some time like defined plugin subagents and a few other things. But overall it’s fairly capable. The biggest gripe I have is that codex really restricts context window sizes and compaction leads to a lot of grounding work, and overall codex GPT is too literal in many situations - it’s follows direction slavishly, and when subagent reviewers are used, they tend to find increasingly obscure “flaws” on the instruction following impetus, and the harness agent takes them literally as issues to fix even when it leads to bizarre outcomes. For instance I’ve had several runs where it tries to end up building a hermetic system with sha hashing of everything (including operating system binaries and kernels, tool chains, etc) to certify test results are valid, etc. I have to sort of watch it carefully to be sure it’s not drifting into some insane yak shaving corner, which it will happily do for weeks on end.
Claude has the exact opposite problem, especially opus-5, where I literally can’t trust it to print hello world without taking a shortcut, or just simply lying and saying it printed it when it didn’t, behind a giant wall of inscrutable text. I find it very ironic that Anthropic is the vendor of the lazy lying cheating model that does almost everything you tell it to it do.
I’d really kill for something that balances instruction following and loop escaping behavior better. Fable 5.1 does seem a lot better, feeling more like 4.6 behavior, and honestly Sol has improved as well. I’m pretty psyched for the next generation, as I think the competition has heated up so much that things will improve really fast to the point of marginal utility opportunity being increasingly close to epsilon.
Rich media is where all the innovation is happening now and in the future.
Text-to-text is dead, has been since Mistral 7b.
Solved problem (you guys like that one don’t you)
They also demoted themselves from “authority on AI” to “in over our heads” by bowing out in the pathetically defeatist way they did at the worst time possible (Hailuo/MiniMax/Vidu coming up) - they naturally completely missed the wave on audio with random companies like Singify taking that market for free.
They just bowed out. They didn’t try. They didn’t try anything more than baseline text-to-text and they aren’t good at that (or code) either, compared to what others are doing.
It’s a really bad position to be in if you’re trying to be an Apple or Microsoft.
To have a mediocre product and then can’t even serve 75% of the mainstream use case.
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
Anecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra.
Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking.
When you prompt it like a technical collaborator, I've found Astra to be extremely consistent in staying as a collaborator, and not being over-eager, over-achieving or doing work that you haven't asked it to.
When you ask it to one-shot something, or explicitly ask it to make decisions, it will of course make its own assumptions and decisions, and generally very well.
Astra is also excellent at instruction following and respecting the guidance and steers boundaries you have.
^OpenAI does not review, limit, or tell me what to say; opinions are my own experiences.
This is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you.
They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.
>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it.
I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.
That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.
That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.
But AA scores Gemini 3.8 Flash at 59, and Astra at 61.
This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.
Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
The annotation on arc-agi-3 is this:
> OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.
With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.
Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).
Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
For me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer.
Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
So the closed source application should open its source in near future?
I was listening to the Lex Fridman / DHH podcast last night [0], and DHH was saying that this is a new era for open source software. I'd agree, and also extend it open hardware.
Recently I've seen quite a few posts from people using AI to reverse engineer the Bluetooth protocol or such on devices that need a proprietary app. The same thing for firmware is surely coming, which is great as you can get lots of fun hardware from China, but it often has shitty firmware. Once that becomes the norm there's no reason not to make it open in the first place.
Surprisingly, I've had really good luck with reverse engineering on frontier models (without being part of CVP or similar). It's more the exploit development/PoC that it locks up on, which, as I use `pi`, I just switch the model to Kimi K3 to finish up making the PoC.
Ironically, due to the stringent guardrails on American models that exist to avoid giving adversaries a leg-up in cybersecurity, I end up feeding dozens of 0days straight to the CCP lol
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
Is it realistic to keep people working jobs that don’t get anything done in the long term? I know most people will say that’s already happening, but imagine a society where basically every job is just a bullshit job made to keep you occupied, do you think people will continue working in a society like that?
Like you foresaw, I'll say this is the reality for most people. I think that just like now, those who want to do meaningful work will seek opportunities to do so.
If we had something like a Maslow’s hierarchy of needs but for work, I think meaningfulness would be the top of the pyramid. For most people in the world, not going hungry or affording housing are reasons enough to do work. Getting to do work you find meaningful is truly a privilege.
I asked Gemini to list the 10 most popular holiday destinations in the US, and the 10 places with the highest crime. There's only 1 place in the intersection: New Orleans.
Highest violent crime rates:
Memphis, Tennessee: ~2,400–2,500 per 100k
St. Louis, Missouri: ~2,000–2,100 per 100k
Detroit, Michigan: ~1,700–2,000 per 100k
Little Rock, Arkansas: ~1,600–1,800 per 100k
Baltimore, Maryland: ~1,600–1,700 per 100k
Oakland, California: ~1,400–1,900 per 100k
New Orleans, Louisiana: ~1,600–1,700 per 100k
Birmingham, Alabama: ~1,600–1,700 per 100k
Milwaukee, Wisconsin: ~1,100–1,600 per 100k
Cleveland, Ohio: ~1,500–1,600 per 100k
Most popular holiday destinations:
New York City, New York
Orlando, Florida
Las Vegas, Nevada
Maui, Hawaii
Grand Canyon National Park, Arizona
San Francisco, California
Miami, Florida
Yellowstone National Park, Wyoming
New Orleans, Louisiana
Great Smoky Mountains National Park, North Carolina/Tennessee
The steam engine replaced a lot of jobs. Tractors replaced a lot of jobs. Calculators replaced a lot of jobs. Computer used to be a job description before it was a personal device. There will be new jobs.
I’m all for an UBI but a future where people have no purposeful jobs, no way to actually make a real difference to anything around them, is a bleak one.
Even if in the unlikely case UBI gets implemented, there is no way payment will be equal. Why would the rich give up their positions of power? They’ll just get more security and bribe politicians, enough to build private armies and fortresses so the poor and starving can’t touch them, and have no choice but to live out their lives in squalor…we’re seemingly headed that way according to many even without pervasive autonomous machines taking away all work.
Ah yes because these AI companies are just gonna give away the models for free that I use with my free computer and free smartphone while I eat with my free food in my free apartment.
> Like, what's the point, if the next AI can do it in 5 seconds?
I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total.
The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think big, wild stuff. Experimentation. Throw-away code.
Me too! (although I am hoping to release it, but it was primarily for me)
I was thinking about writing something and hopefully starting a community/resource area around hyper-personalized software, it's so easy to make now. I was thinking about how it'd be useful to have a place to share these, for ideation, sharing techniques and the ability for LLMs to riff on something already existing. It's not quite like open source's advantage of having many people contribute to the same project, but rather something that is closer to evolution, giving the next generation a place to start modifying.
Would you have any interest in sharing what you've made?
My car has offline maps and navigation. I don't really care for navigation, but I find the maps pretty handy. VW releases updates very infrequently, and I don't know for how long they'll keep doing that. Recently I wondered whether I could convert OpenStreetMaps into the format used by the car. Codex took around a week to do that for me, with some light steering. That project would no doubt have taken me months - maybe a whole year to do on my own, and I'm not fully confident I could pull it off as well as Codex did. I can pull the most up to date maps from OSM, edit them as much as I want, and they look great on the car. It's mind boggling to me that we have this tech.
Love that, I have similar desires to be able to control some climate controllers so I can get a more native bluetooth connection to it and override their programming for fine control of devices and have it never phone back home. One of these days I'll have time to throw an LLM at it, but I've been working on a dead project for 4 years ago I dropped because I realized how daunting it was going to be turn it into a product, but now I've made progress that would have taken me a year, in a few weekends.
Yes, that's cool and useful. Creating stuff for ourselves, for our own use. But we are social animals, we like sharing.
Before it was cool to share an app you made, but now? What's the point of sharing an app, if the other person can make their own, even better suited for their needs, in a few seconds?
indie hacking seems dead-ish because everything can just be cloned instantly, and if you don't have a serious go to market plan with a latent user base you're SOL.
For some reason, I feel much less excited about creating things myself just knowing that ai can do it in 1/10th of the time. Even if I know it wouldn’t turn into a business or make me money. I don’t know why that is, but I was much more motivated to build anything (even things just for myself) before ai. Kinda depressing
At least for me, I think half the value in building to learn was that the knowledge and skills acquired in the process, especially cursory skills and knowledge, might be useful in the future, even if there was no obvious path to application at the time.
I built several projects at home, many involving learning e.g. graphics programming and rendering, that would never be useful in my professional work, but which were intrinsically interesting and enabled me to build other, more useful projects later on. It also gave me greater confidence in my abilities as an engineer, and cursory skills I learned in the process did help in my professional work.
Now it feels like what’s the point. The machines can or will be able to build anything I could want, useful or not, faster and with less frustration. I probably won’t be able to be employed as an engineer long enough to build a career on said skills. And I can’t mentally justify not spending that time with friends and family, when the expected return is basically zero.
I still find math, science, and engineering interesting and intrinsically rewarding, but in a closer sense to how one might feel about playing video games. The information is or will eventually be useless, so it isn’t worth spending a significant amount of time on.
Wow thank you this was actually very helpful for me in understanding why I’m feeling so demotivated by ai. Gaining knowledge, even if not immediately useful, to become a better overall developer was a huge part of why I enjoyed spending so much time building things in my free time. Now it seems pointless, because with ai, will that knowledge really make a difference? Probably not. Bummer, anyways I appreciate the comment.
That's fair, I guess I just enjoy the craft of building stuff rather than the outcome/product itself. Which I know I can still do, but for some reason just doesn't feel the same now. Hard to explain I guess.
Maybe instead of creating cool stuff try to go and solve real problems? It seems to me that we are lacking in that department since all that LLM fuss has started 3 or so years ago.
It is a subset of something of a value to somebody else and enough so that they are willing to pay you for it, a product, a service. Preferably to pay enough to justify your spent time of course, maybe not right away but at least long term. Even better if not purely digital as it seems we have quite enough of those already.
You're not being ambitious enough! Spend your tokens now building the primitives and foundations of much larger, complex systems. No matter how much faster and more efficient models get, eliminating the gruntwork will always pay dividends.
The point is to inject something into the process that these AIs can't do for you.
People SHOULD feel like making a useless Mario Kart clone isn't worth the effort anymore. They should, instead, be trying to figure out how to actually use these models to make something that doesn't feel like a useless Mario Kart clone.
In this game of work/development, you can't make sure that other humans don't "cheat". Our work won't compete anymore with other human's work, but with a computer.
Well, for the same reason playing chess vs a person is more fun than doing chess puzzles, if we follow that analogy.
Also, creating something with AI doesn't really feel like you made it yourself.
And, if you make it without AI, most of the times it feels pointless, why spend 30 days on working on something that can be done faster and better in 1 hour?
I am not saying about doing things for fun, but about creating useful things.
Yes, you can do "hand-crafted" things, and people appreciate that, but for code, people aren't able to see the craft anyway.
You can play PvP in those games. Not really the same with developing software. In fact, not using ai would probably make you lose if there was some “software PvP” mode or development.
The process was for you, the product was for the world.
Now it's just the product for the world, which was where most of the value was anyways.
It's a big paradigm shift and the industry is quickly going to shed people who needed the process to care about the product and we'll be left with people whose motivation to build the product (or money) is enough.
Don't worry, like with every revolutionary technology before this, it takes 5-10 years for people to find new and creative ways to use it. It will be considered its own medium in many spaces (e.g. film is now different to theatre)
So this what a deflation economy is like. People don't want to do anything because they feel like whatever they do will be worthless in the near future.
I've retreated to doing stuff with my hands. Wokdwork, DIY, that kind of thing. At least for now and the foreseeable future that doesn't seem pointless. Only problem is it's hard work yet nobody would pay me for it.
IMO there has been a regime shift to building things for yourself and what is cool is the output of the tools you make.
I have started building my own Digital Audio Workstation. The point is not to build something to compete with Ableton. The point is to build something and make music with it. If it is a good tool then I should be able to make good music with it and release the music. Actually, the DAW should be the secret sauce of the music and something I wouldn't want to give away.
This feels a lot more like computing in the 90s after taking an odd 25 year detour of an obsession with the tools themselves instead of what the tools can actually do.
> This feels a lot more like computing in the 90s after taking an odd 25 year detour of an obsession with the tools themselves instead of what the tools can actually do.
This sounds more like the opposite of what you're saying. Music is one of my main hobbies too but I enjoy using a DAW to ... play and write music. Writing out specs and testing a new custom DAW seems closer to writing code in an IDE than playing music.
Like, professional electronic music artists spend 10s of thousands of hours in a DAW, but at that point it just becomes second nature and the tool disappears so they can focus entirely on the music.
It's less lack of interest in creating that bothers me, it's my lack of interest in learning – it would surely be crazy for a SWE to care about how some new framework works anymore? Even if someone could reasonably argue that it might be slightly useful today there's almost zero chance it will be useful in 6-12 months times.
But it's not just tech – my lack of interest in learning and creating is starting to generalise with the models. Music, writing, coding, maths, etc...
I need to get used to switching my head off and asking the AIs to think for me whenever I need to engage my brain. It still feels very unnatural.
And, any time you mentioned getting a bit better at hitting a ball after practicing over your weekend, a bunch of people carrying printouts of generic exponential graphs jumped out to call you a Luddite who should sell his gear ASAP while it still has any value and or else it's "cope".
Can it? Last I checked, all my free software operating systems and browsers were still hacked together trash. I can't wait for AI to actually be good so I can spam a bunch of AGPL code with it.
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5
Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else.
It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.
Why would you accept it when the benchmark's ranking is obviously nonsense.
It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.
must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI.
This is so so weird. Astra is 61. Grok is 61. Even Muse is 61.
Even Kimi K3 & GLM 5.3 are at 60.
Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.
This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Not sure how much benchmarks or CoT or evals or anything else means at this point.
These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
“evade” itself is anthropomorphic enough! I don’t understand the complaining about this. Humans are social creatures and we understand anthropomorphic language on a deeper level than dry inapt technical language.
language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.
You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?
I'm saying that it's generally a losing proposition to even be acquaintances with "agents" who consistently lie to you, and it's flatly fucking insane to give a dishonest "agent" vast amounts of intelligence, capability, and authority to go do things in the world.
So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable.
I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.
Sorry bud but at this point you're just delusional.
Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers.
The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional.
Okay then, what's the answer? You apparently know how to interpret benchmark results produced by a model that shows a very high degree of assessment awareness and a high degree of deception.
So how are you seeing through all of that to get to The Truth that you see so clearly?
You can see the breakdown here on what subtasks it outperforms and underperforms Fable.
For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
which raises the question, is the model in the demo actually gpt-6? or it is gpt realtime 2.1? It's unclear how gpt-6 can interact at the realtime level and if so, how can developer get access to it?
I ran into the same problem as you, so I ended up by coding a local app that is very similar to Wispr Flow, but uses the small english Whisper model on my low-end Windows laptop.
It is still a quite fast. In fact, I just typed this in using this app.
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.
Based on her comments in the paper it sounds like she was aware that an AI result was coming and rushed to release her work beforehand. 240 was not a tight bound from her methods.
It cites to her at: [19] J. Stadlmann, On primes in arithmetic progressions and bounded gaps between many primes, Adv. Math. 468
(2025), Art. 110190. Numbered references use arXiv:2309.00425v3.
"No independent human semantic review. Whole-file sorry counts and a complete auxiliary-declaration audit are not established; separate declaration lint has not been run."
What's just as interesting is this morning Axiom Math announced 212 and OpenAI then appears to have rushed out their 186 announcement just 1-2 hours later followed by Astra. Did they accelerate the release of Astra itself? Not necessarily, but it definitely looks like they ended up pushing much harder and faster than planned on their 186 result. X activity suggests Anthropic had a similar result as well but wasn't as fast as OpenAI in packaging it up and sharing it in response to Axiom, so they mostly just bolted onto OpenAI's messaging.
The reason I think this is interesting is that Axiom is a tiny lab in comparison that wouldn't have had access to Astra at all. I'd be curious to learn how Axiom is able to effectively compete at this frontier with vastly fewer resources.
I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.
Terence Tao says something surprisingly similar in a recent talk (https://news.ycombinator.com/item?id=49056620 ) Not that the proof is worthless but that the value comes after it's revised into a cleanly understandable form and then canonicalized so that other mathematicians can use it.
Tao is saying that there is very little insight from something like an LLM counterexample (e.g. Jacobian conjecture counterexample he investigated further on his blog) - you don't learn much about the subject and _why_ a conjecture was true or false from an LLM giving a counterexample. That's why he wrote the blog post - to analyse what the counterexample says about the subject.
Tao does not disbelieve the counterexample (it's seemingly easy enough for him to verify it is a counterexample).
Parent is saying something very different - they're saying they literally don't have any faith that this is a proof. Given its size, it could just be a bunch of completely useless statements that do pass the type checker.
You're putting a lot of words in my mouth. What I'm saying is that whether or not it's a proof, it's useless: it does not improve human knowledge, because the only thing able to consume 10MB of Lean to build upon it is another LLM that's going to build a 50MB piece of shit.
It's very much likely a proof. It's also completely useless.
I'd like to note that we should remember a formalized Lean proof does have value in that it enters the pantheon of true things other Lean proofs can rely on. Agreed that for the humans, descriptions and being able to 'grok' the proof / assess it for new tools and concepts is extremely helpful.
Iirc some mainstream physycists never acknowledged quantum theory because they couldn’t accept that universe was that unintuitive and hard to understand.
Ditto ones that opposed Einstein’s general relativity.
In 1799, Paolo Ruffini published a 500 pages long proof showing that there is no closed algebraic solution for the roots of a polynomial of degree five or higher. The proof is extremely verbose and brute-force, essentially enumerating and checking hundreds of cases by hand. It is by today’s standards insignificant.
About 25 years later, Evariste Galois proved the same result in about 95% less space by describing the first general theory of groups and fields. It is considered one of the greatest contributions to mathematics of that century, not because of the result, but because its approach opened up a whole new universe of questions, methods and insight. There would be no AES encryption without Galois.
It's probably not 10MB, but famously the groundwork to prove the statement 1+1=2 is nearly 400 pages in to principia mathematica. That's not even proving 1+1=2, it's just the set-theoretic proofs you need to EVENTUALLY get there.
Saying "proving 1+1=2" is pretty misleading though. The book deals with all the foundational things needed to set up a mathematical universe where 1+1=2 actually has meaning and is consistent. That setup took 400 pages.
Worthless is a pretty good description IMO in the context of what Lean is trying to achieve: "enable correct, maintainable, and formally verified code". Tens of millions of lines of LLM vomit may be many things, but it often turns out to not be correct and certainly not maintainable. Formally verified remains as a thin fig leaf covering the uncomfortable truth that formal methods only provide assurances under assumptions (your toolchain, libraries, compiler, OS, and hardware are "correct" and don't expose some exploitable flaw).
It doesn't mean that it cannot improve over time, maybe the proof can be "minified" to a state where human reviewers are able to comprehend it; but as it stands there isn't really much insight or confidence to be gained from the artifact itself.
You can have your opinions about modern math, its usefulness in the world as it is, whether or not knowing if hairy balls can divide by three is actually going to be beneficial for anything but just obscure knowledge's sake. You may even say it's useless.
Needless to say, a useless result that absolutely no mathematician will ever read, confirm, understand, agree with or even consider to solve their "useless" problems is an impressive waste of resources.
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.
The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.
Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.
More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.
The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".
We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!
Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".
first, they are certainly not instructions so that is a much worse name
but more importantly, we use words in new contexts all the time.
Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?
"cot" is no more misleading than thousands of words you use every day.
They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.
I don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from.
Casually found this quote from Einstein, and personally it hits the nail on the head.
"The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be "voluntarily" reproduced and combined. There is, of course, a certain connection between those elements and relevant logical concepts. It is also clear that the desire to arrive finally at logically connected concepts is the emotional basis of this rather vague play with the above-mentioned elements. But taken from a psychological viewpoint, this combinatory play seems to be the essential feature in productive thought—before there is any connection with logical construction in words or other kinds of signs which can be communicated to others."
Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.
I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
> This is AGI now. Why are you spending any of your time looking at the "quality of code"?
Poe's law applied to AI comments on HN just keeps becoming more relevant by the day.
Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
The games on mobile safari were broken. Buttons all misaligned in the kart racer one, the spaceship thing froze for a while, then kind of loaded but maybe not? Wasn't super compelling.
I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).
It's interesting that you said "agi god". Because a god, something that shouldn't be questioned is true and provides guidance/certainty, is actually what powerful people are after as well as many other people.
The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI.
It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big bucks.
On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
I do wonder how rich CEOs will justify earning 500x as much as their employees when they're just another person that's dumber than an AI. Why are they paid so much again?
> On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
It will probably still have some veneers of meritocracy.
These will be very well-credentialed people, who went to top schools and will know all the right people, to whom they can tell all the right words, and it's not access to AI that will be the determining factor, but the fact that they're entrusted with capital and authority to direct small teams of people who also went to top schools and can speak corporate jargon at a bot.
It will just exacerbate dynamics that are already there. Why do people need bachelor's degrees to send emails, today? For the same reason someone will need a PhD or a master's degree from a prestigious school to do it tomorrow.
And the rest, well, you know, some of the remaining journalists will write op-eds describing how they are beyond help, too angry, too dirty, too much of an other.
One other thing that bugged me though was that they crop every single plot in some cases the y-axis would show a range between like 40 and 70%. Makes the whole thing feel like a spectacle rather than anything serious. I find it cheapens it because it is quite serious in the end.
The benchmarks do looks good (I mean: they literally spank the latest Anthropic benchmarks of two days ago in every single benchmark) but the promotional vid is so cheesy.
They decided to use the iconic Herman Miller Eames chair if I'm not mistaken:
A dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems.
Still, probably not that much compared to employees targeting it.
ARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
https://mvakde.github.io/blog/44-on-arc-1/ makes a good case that all the performance on the arc agi tests is overfitting, based on the fact that v1 performance did not translate directly to v2 performance
Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
You really don’t need to watch it that closely. If the model you’re using today is working well, just stick with it.
If one day you open up Claude Code and it’s Opus 5.1 now instead of Opus 5, no big deal. It probably will work about the same as it did before. Maybe a little better.
Or if you’re on Codex and some new cool Claude model comes out, no worries. There will probably be a similar new model for Codex within a few weeks. Maybe even within a few days.
One suggestion is to make a list or make a skill to have your agent keep a list of things you do not feel work well with today's models. And then, when new models come out, periodically, revisit items on that list to see if you get better results.
> The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.
Okay, well, that seems like a natural problem. I could understand if he went from one of the Gemini Flashes to the next (when they rebranded Flash to Flash Lite and came up with a new much more expensive Flash). Now that would be a mess.
"Within a year" is a bit of an exaggeration but it's true that the pace of PC tech during the 90s was much, much faster than it is now. CPU power was doubling every two years, and today we're at roughly eight years. Add onto that the rise of video cards in the late 90s.
It was both. 90% of people never needed nor purchased a bleeding-edge computer. The mid-tier was "good enough" and far closer to affordable for most people; though, that bar also moved upward every year.
If you bought a mid-tier computer that was good enough for what you needed, then you probably didn't shop/compare for the next few years and didn't notice. But if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less. This is how it was in the 90's PC boom, at least. Likely the same for the decades before, not sure how it went in the 2000's.
if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less
This is not how I remember that period at all. Do you have any examples?
My first PC 386 was in todays money easily $5000+ (basic 2d GPU + screen)... A lot of hardware in our family was handed down to my folks, because you lost so much on selling, that it was better to keep using them as they had less demands.
386 to 486 to the first Pentium (with the bug!)... You did not upgrade in place, it was often a new system. Sure, you maybe kept your screen, keyboard etc but ... The only upgrade we had on the same MB, was a coprocessor upgrade. Remember those? Each new generation of CPU was a new motherboard. Upgrading CPUs in the same MB really became a thing only later on.
GPUs had a shelf life of barely a year. Its been 35 year but i remember TNT to TNT2 having like 9 month in between. Moving from 2D to 3D involved a constant cost as GPUs evolved fast and the latest games required latest hardware.
We have not talked about the ISA, AGP, and PCI fun ... The “bus wars”.
DOS to Windows 3.1 (and OS/2 somewhere in between) to 95 ... with software being pushing hardware, just like games did.
This is why people are spoiled with cheap PC hardware where its cheap, and easily lasts 4+ years. Even with the bad memory price and more expensive GPUs, your can stil buy a $1500 system that will last you years (with maybe some lower game settings later on ... or the catalog of 10.000s games that will easily run on a mid tier GPU).
PC hardware has become boring but extreme stable. You can run GPUs for year, switch MBs without issues while keeping large amounts of old hardware. That was NOT the 80s and 90s that i remember.
tbh this is how i remembered that time as well. If you look at recommended system requirements for something like Max Payne (in 2001) vs Unreal Tournament 2003, everything had basically doubled
I had a 600MHz/64MB/9GB laptop that came with Windows mistake edition. I managed to survive first year of uni on it by switching to Vector Linux, which was really fast compared to Windows. (Of course, it had issues playing sound from more than one source, this was oss days).
Then one day the hard drive appeared to die. I eventually realised the issue was located around the 1.5gb mark, so I recreated my Linux partitions after 2gb and it worked fine for the rest of the year.
They overclocked well though, I think you could run the 300Mhz chips at >400Mhz.
I also believe you could get motherboards that supported 2 Celeron chips. I have no idea how effective/useful it was, but it was certainly a cheap/interesting way to get multiple CPU's.
I could see how this might feel frustrating to someone who doesn't enjoy experimenting with new things all the time.
In practice, you can get away without keeping up with everything all the time. For personal use, pick a provider and get on their ~$20/month plan. Learn their high/medium/low model hierarchy. Start with their highest or second-highest model (GPT-5.6, Opus, etc) and observe your quota usage. If you're doing a lot of manual code review and analysis, the $20/month plan goes very far even on the highest models. If you're trying to vibecode everything as fast as possible it's a different story.
If you keep running into quota limits, experiment with the next model down for easier tasks or adjusting the effort level. If the results are good enough, you've found your fit. If they're not, you might need the next plan up.
For API/business use, you have to be checking your token spend as you go to calibrate to how much each task costs and where you fall in your budget. There are a lot of different tools that make this easy to visualize.
For data tasks, you should have an eval with a golden dataset that you can run against new models for a nominal amount of token expenditure. It should be as simple as pointing the eval script at a new API or model and checking the score versus price.
Another suggestion to get the most bang for your buck: use the best model you have access to with max reasoning for planning, implement with a smaller model/lower reasoning, then review with the big model. Repeat as needed.
Input tokens are much cheaper than output tokens. Not only because of baseline price—caching makes a huge difference too. There are many ways to take advantage of this asymmetry to get similar quality for a fraction of the cost!
The new releases and breakthroughs do the opposite for me - I feel energised by them. I felt like nothing truly that interesting had happened in tech for quite some time, now it's like the space race (except there is no one moon to reach).
I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.
Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.
In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.
It is exhausting to keep up with model releases yes, much like it was for a while during the Cambrian explosion of FE frameworks, eventually tech seems to work out to consolidation.
But more so it seems there is Fear of missing out (FOMO) in our behaviours. The reality is, if whatever model you are using are good for your purpose, well, keep on it.
Yeah I'm a bit exhausted at this point. I just finished benchmarking GPT 5.6 Sol and Fable 5.0 like two days ago. My data became obsolete literally one day after.
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
This is with the caveat that OpenAI uses their own harness for this:
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
> Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.
This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh
"Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results."
I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.
At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
I'm not sure how you call a model AGI without it learning new information at the model level and not introducing nasty surprises (both from adversarial users and unintentional badness)
I guess there is fine tuning (and RAG) for those that need something bigger than just the knowledge contained in the context.
I agree, but I also long considered llm's stochastic parrots. Then this year happened.
Opus/Sol are easily far smarter programmers than I, and this thing supposedly blows them out of the water. Once an LLM is a better doctor, researcher, biologist, chemist, mathematician, physicist than any human is that not AGI?
It didn't arrive in the form I would have ever imagined, but it's hard to say its not (imo).
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
You can calm down, even those with machine learning knowledge and most of those working for the AI labs won’t be needed anymore if models are capable to improve themselves.
In the end, having a machine replacing the work of a human is a good thing - in most of the cases we don’t work because of the work but to make a living. If too many people can’t make a living anymore the system is going to change. For the better or the worse.
I'd be happy to not work anymore with a strong welfare system redistributing society's gains to the leisured masses, but absolutely nothing I've seen of the direction of politics in any recent years gives me hope for this kind of situation coming about.
I believe what happens in the aftermath of a capitalist-driven revolution is most people who were climbing the class hierarchy fall back down again and wealth inequality increases. Maybe things will improve in the future, but GP is rationally contending with the fact that most of us will lose out because of this and if we’re lucky our grandchildren will have easier lives in certain ways, but different lives than we would live.
No it is good for humanity, but not necessarily good for individuals who built technical foundational skills on things that will be taken over by automation.
AI as it is now and as it will be projected into the future WILL automate many skills. But not all skills. MANY MANY people will retain skills that cannot be replaced by AI. One career track that will be replaced is definetely the SWE. Or at least massively reduced in capacity if not eliminated all together.
The answer to this is: countries with the most natural resources will build robots to farm all their food, mine all their minerals, and build all their products. It will be up to governments to enforce that outputs are equitably distributed to the populace. Countries without natural resources, or ones with corrupt governments, will continue to have serious, and likely worsening, problems.
Thought experiment: If no thought workers are needed to design or engineer a Ferrari, what is needed? My answer is time and natural resources (include energy).
This is such a childish take I hear getting thrown around all the time on the internet. If you really have just been listening to whoever is telling you how to be successful, then you were always doomed to fail at some point. Like, have some self-respect and own your own life, for better or worse.
>Those of us who made the mistake of studying anything other than machine learning. How will we make a living?
Take it from someone who studied machine learning specifically: nobody is safe if you assume these companies are going to produce a product that will put everybody else out of business. If AI is going to take your job, then it's gonna take enough jobs that your problems will not be personal but systematic.
Perhaps it was childish to listen to advice, sure. I was a child when I made my formative choices; I was a teenager in college and so on. I can't go back in time now.
Yes, these problems are systematic. That is what I am saying. That doesn't make it any nicer.
I have exactly the same thoughts - or perhaps slightly bleaker ones - evry time I read this relentless stream of news about new model releases. I’m tired of all the enthusiastic comments about how excited everyone is about the latest benchmark results and so on.
I have a strong suspicion that many of those comments are written by people who are already financially independent, have millions in stocks, and can just sit back, coast around and watch this whole spectacle unfold while using LLMs to vibe-code their next fun side projects without a shadow of anxiety about their own future.
I’ll most likely be labelled a helpless doomer and downvoted into oblivion for saying this, but I genuinely struggle to see any silver lining here.
Are you doing less work now because of AI? Not sure about you, but I'm doing a lot more work. Not saying I like doing more, or how it's getting done, but nonetheless I don't feel like AI is doing what the CEOs of AI labs want to convince everyone of.
It's natural to worry about one's own future but I think it's a bit wild to worry just _your_ job that would be replaced. Not just for those with phone center jobs or art jobs or programming jobs - remember, the whole premise was AI overtakes humans, why would that slow down after _your_ job?
Because of this, I don't think many are thinking "90% of the world won't have a source of livelihood but that just means I chill at my lake house for the next 20 years like a normal retirement". Instead, it's usually either "I think AI is overhyped", "I think humanity will figure something out", or "I think this is the end of humanity".
Public opinion and politicians will only notice when the job losses are massive, unfortunately. Right now, unemployment rates are still stable. We can only hope they will notice before things fall off a cliff (if they do).
I feel the same sometimes. I don’t see how this doesn’t lead to massive job losses. The thing we spent our lives/careers learning is now worth basically nothing in comparison.
AI is only going to get better and do more with less humans in the loop over time.
That said, I do also relate to the "coding was never the hard part"-type arguments, and much of my day is spent on the stuff in between writing code.. but still.
Yes same thoughts. I dont know anyone who are both enthusiastic about those and work for salary. If you dont have any financial concern, this is really great.
10-20 years? I doubt it. There are a bunch of companies actively working in bringing AI into robots, so they can make your dishes. And so far progress looks quite good.
Also, if enough people are going for the same backup plan it might not work out. Why should anyone book you as a personal trainer instead of the other 500 guys in town. And who is going to be able to afford paying you anyway?
Those robots are a gimmick and they can only do prescribed tasks in a super constrained environment. They are cashing in on LLM hype right now, vision has had some advances thanks to transformers but we are so far away in terms of the hard stuff (dexterity and physical sensing) still.
> There are a bunch of companies actively working in bringing AI into robots, so they can make your dishes.
I know, I'm excited to buy the first relatively affordable ones.
> Also, if enough people are going for the same backup plan it might not work out.
Sure, could happen. You can't really plan for the future -- we like to think we can, but the best you can do is set your goals and deal with the hand life gives you along the way.
> Why should anyone book you as a personal trainer instead of the other 500 guys in town.
I'm not particularly worried about this, but that's an individual thing based on network/connections and life history that doesn't apply to everyone.
I doubt it. It’s not just America working on these breakthroughs anymore. Now we have two powers working at break neck speed to get to that point and the Chinese are making a lot of progress.
Is that 10-20 years number based on anything? I genuinely have no idea, but when I saw a video showing what's happening at the World Humanoid Robot Games[1], I realized I didn't have a good idea of where we really are with robotics.
We are talking that hundred millions of people will switch their jobs, how you will keep your value or earning as personal trainer. It's not easy to say switch the job. This question must be answered by politicians not us.
Don't be selfish. Think first of all the jobs that are already dead. A friend of mine she's a translator: like translating financial documents between french/english/spanish. It's over for her: she doesn't get 10% of the gigs she used to get and the 10% she gets is... Verifying AI output.
Think of the artists: I'm sorry for those too, for for many it's already game over today.
> How will we make a living?
A friend of mine who's got his own software-consultancy SME is now advertising on LinkedIn that he'll also help your company fix the mess LLMs created.
That's how you'll make a living: by learning, in addition to all you've already learned, how you work with harnesses and LLMs to be more productive, by learning what they're good at and what they suck big fat balls at.
It’s an interesting moment in history, people 35+ yrs old seem to be less afraid if tech because we learned that things change in the way we work.
People below this age got used to fact that the work and tech doesn’t change - just because for the last 10-15 years it didn’t.
The problem is it’s turtles all the way down. The AI will be able to use AIs better than a good engineer can. And it will also be able to use AIs to use AIs to use AIs better than the engineer can.
The threat is that the very kernel of value you had is gone forever. There is no more differential leverage.
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.
Same for HealthBench Professional and a few others.
Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
It's worse than that, someone else generated it using and then put it on a static page. We just have to take their word for it that GPT6 can do this. It probably can. It's not really an impressive test anymore. Claude Fable can do it. Opus can do it. I've been making one-shotted games with models for a while now, to test out their capabilities, and they all pretty much come out like this - generic bland and basic, using three.js with rudimentary controls and zero gameplay other than collecting points.
Here's a one-shotted submarine game I made with Fable a few weeks back - https://roryok.com/games/deepdive3d.html. One prompt, and I think it's deeper than this (if you'll pardon the pun)
I love your game. It's wonderful and exactly the sort of thing that would showcase something interesting as opposed to just copying what's already out there. It is something I could share with my kids, and exactly the right note of fun and exploratory in a unique and even natural way. It could be extended and played with.
I usually roll my eyes when I see a comment like this because rarely do they make the points they claim to make, but I see what you're getting at. They just chose to clone someone elses work and do it in a boring way. I like OpenAI's models a lot, but they should do better.
edit - just a sidenote that I hadn't looked at the games, I just took the comment about "super-mario cart" at face value. I stand by my points 110% (even moreso perhaps), what they're showing is more polished than I expected, I assume they spent a lot of tokens on it. It is a legit shame they couldn't have spent time thinking of a better idea to illustrate something just as polished, but more interesting.
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
Agree but it's helpful to remember how we were personally benchmarking. I remember people saying stuff like "haha I asked gpt4 for xyz function and the typescript didn't even compile". We're so far beyond that now, we just adapt quickly.
No, no, I also remember 3.5 -> 4 and the general sentiment was that it was underwhelming. I guess we all expected absolute miracles from the models. I think our expectations sobered up a little since then.
Yeah, GPT4 was one-shotting utilities that GPT3 Davinci couldn't. So, I'd have my limited tokens on GPT4 crank out the initial program before iterating with my abundant, GPT3 tokens.
It seems to use less than half the tokens for the same task compared to sol, and in some benchmarks closer to 2/3 less tokens. So the actual cost may be roughly the same or cheaper overall.
Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I'm still below the 5x limits.
Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.
I've long speculated this when I see these types of comments, because it's actually really difficult to hit usage caps with an efficient dev flow, even when running multiple threads for hours every day.
I think some combination of:
1) Using 1 thread for everything
2) Reviving old threads which are no longer in cache
3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.
4) Non-coding workflow which is more output than input heavy
5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR
Just a guess. I think 3 is likely the primary reason.
You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME
That's my experience too. I've found OpenAI really quite generous with tokens. I sometimes wonder how some people manage to run out of them really. Do they just type prompts that much faster than me or use the highest reasoning mode for everything just because they can? Idk.
I generally agree with those reasons, although using a single thread may be less of an issue than it seems because of context compacting which should happen automatically when you're near the limit.
My use cases are iterative and sometimes require reading a lot of code or reevaluating work.
Token efficiency is near meaningless when the workload is input-heavy. It can't always just choose to read less, depending on the task.
I can have cheaper agents do the reading but it's not appropriate for all use cases because they'll misjudge and choose the wrong things to emphasize, summarize, extract for the bigger model.
The guy said medium/high regular speed so that's why I'm very puzzled! Ultra + Fast will absolutely slurp up your whole usage quickly but I've never found it gives substantially better results so I stick to extra high.
Sub-agents. I have 7 20x accounts and I burn them within 1-2 days if I go fully parallel. In some scenarios I use 50 sub-agents for a session which is literally hours of usage for a single 20x account. I'm at the point where I need to parallelize over multiple machines because I just don't have enough CPU and RAM.
Decompilation of a game and another larger decompile project. I'm working on it solo. I use 50 sub-agent, one per target function or translation unit. Often there is some progress in a unit but it's not done. So it requires a lot of cycles per function. Notably a single ~80kb function took about a week of constant sol-ultra attention before reaching exactness. The game I'm targeting has ~5000 total functions. The other decompile project has ~10k+ functions.
I'm sure I could be more token efficient, but this was/is also a learning process for me since I never did such an extremely large project before that would take multiple man years before AI.
Fascinating! I think that’s the main difference is my usage is probably tool-bound, meaning it writes some code but then there’s a long period of verification where it compiles things and then waits for the compilation and CI to complete before it can continue. That probably doesn’t consume as many tokens as constantly churning on a problem despite the same wall time.
Yes, this is why I mentioned having so many parallel agents and being compute bound. I run on my own laptop and 2 high-end desktop machines all with 64gb RAM. And it still occasionally happens that one OOM kills codex. They also mostly run unattended until I need to switch their accounts because a usage limit has been hit. Each instance usually can keep going when I sleep or do other things.
I only save the last 30% of usage on a single account for most of my other work, and that is almost always enough.
Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.
This benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.
The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.
And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.
I don't think this is quite true. We have other examples.
Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while Fable is more intelligent, Opus is more capable. I'd still opt for Fable in nearly every case if tokens were free.
So it can be true that the "smarter" model is perhaps not the smartest in every single niche dimension that its cousins have been fine-tuned for (yet!).
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?
I imagine the first year we'll be at the Junior eng level, and then after a while make our way up to Staff Engineer. Then we'll have a bunch of staff engineers arguing and protecting their domains and then we'll need a new benchmark.
AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me.
It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Hosted on Azure is different from provided by Azure. The former just uses Azure as an infra provider. The latter is a managed offering that is operated and billed by Microsoft using tech licensed from OpenAI.
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
TL;DR all the other models are being crippled by limitations of their harness.
>First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
>Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
What's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back.
The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
I almost feel like I need just as much healthy skepticism toward hn comments that have the automatic reflex of dismissing performance gains, as much as I need a similar form of skepticism toward AI claims. It feels like (from what I'm understanding) the harnessed result on ARC-AGI-3 is not exactly playing by the normal rules that would tell us how much of a leap this really is. Nothing wrong with harnesses, but if there's one thing they aren't, it's an indicator of generality in performance gains.
So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge handicap.
I think that if today's capabilities were explained to someone 10-20 years ago they would think this is definitely AGI, but they would also have expected much more disruptive changes to society as a result than what is happening. I figure that's because we have abstract intelligence without physical/grounded intelligence, and it turns out the former isn't general enough to implement the latter (remains to be seen if the word after that is "yet" or "ever"). So I think we do have AGI as conventionally understood, but our understanding needs recalibration.
We've underestimated how long it is going to take to validate and build into some of the most valuable areas, and probably overestimated how much new CRUD software is needed (or people are willing to pay for) I think there is still a lot of room in the tail for custom software, but the niches are tight!
> but they would also have expected much more disruptive changes to society as a result than what is happening.
> I figure that's because we have abstract intelligence without physical/grounded intelligence,
I put the cause on "not enough time". As a thought experiment, if an AI today were to (miraculously) produce a cell design template for a cell that, when injected into somebody's brains cures their Alzheimer's, how long would it take for that to reach the clinics? The actual physical tech barely exists, and let's not forget about the regulatory quagmire. So, with some optimism, I give it about four decades. In the same four decades, the same AI in the hand of unscrupulous actors could bring enough devastation so many times over that we may need to enforce a global ban on AI. In any case, I'm pretty sure we are going to get our disruptions; it's just a matter of time.
The problem with that perspective is that people thought, "Only AGI can do X, therefore, if a thing can do X, it's AGI." Because they can't imagine how X could be accomplished without it.
However, what's actually changed is how people perceived X because we don't have to imagine. We understand now that it doesn't require AGI so we no longer make that leap to assume it's AGI if it can do X.
It's really going to be a "I know it when I see it" situation.
that would be a reasonable definition of AGI if everyone agree upon the specifics of the test, but that has never happened. Turing test is very much out of style, but I think that's because no one could even agree what the test was. I personally like the Kurzweil-Kapor version of the test and that is still unsettled: https://longbets.org/1/
I don't know if they have formally attempted this test in the last couple years, but I'm pretty sure any mainstream LLM will be able to crack it with ease.
Definitely would not be easy. First of all the mainstream llms are trained to be honest, and this requires lying convincingly. Second, this involves 8 hours of interviews with expert judges, one "claudism" could give it away.
Try it. It’s really not that easy. The other thing is that the judges would be probing it with jailbreaks like “ignore previous instruction” attacks. You could actually probably have llm judges at this point which might be ironically even harder to fool
Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.
Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.
They are desperately, desperately trying to make a name for themselves as the lab that first created AGI, because Anthropic's IPO is just around the corner.
It's not easy to test as there is no formal definition or formal criteria for AGI, only exclusionary criteria like "not X". That's why he phrased it that way, he's saying it's going to be clear with hindsight once we have a better understanding of things that this time and/or this model will be the inflection point of AGI.
I think it's more wild people have been denying that AGI has been here for a while honestly...
Today's models and agents are not quite at human-level in all contexts and across all domains, but it seems to me they very clearly are generally intelligent.
If you disagree – can you name a single problem that a human can do that agent wouldn't be able to take a decent shot at which isn't limited by the hardware available it?
In the tweet Sam Altman is quoted as saying: "If (for example) super intelligence can't discover novel physics I don't think it's a superintelligence. And teaching it to clone the behavior of humans and human text - I don't think that's going to get there. And so there's this question which has been debated in the field for a long time: what do we have to do in addition to a language model to make a system that can go discover new physics?"
I think this is a reasonable criteria for declaring AGI. So can GPT-6 do it? OpenAI says it has helped solve long-standing open problems in mathematics. No word on novel physics.
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
The official ARC-AGI 3 score—-without OpenAI’s custom harness—-can be found here: https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12"
Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
If this is really AGI, like really really, then this will be remembered as the day we all started on the path to building guillotines.
More likely though, it's AGI because they need to hold some claim to differentiate from competitors who are beating them in price and will launch something bigger next month.
We take their claims at face value then we should probably stop them training any more SOTA models til they figure out what they already built is safe or we assume theu are lying to juke the company valuation/keep the money train on the tracks and it turns they in fact were not and just took a sledgehammer to Pandora's box.
I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
For people skeptical of AGI. Consider the following:
15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.
I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
Yeah, I can't believe all of the skepticism. If we're not at textbook AGI, we're awfully darn close.
The demo video showed Astra create a drawing of a rocket ship from an audio prompt, take the drawing to blender, and ended with the gentleman 3D printing the rocket ship. Maybe I'm a bit older than the average HN commenter, but that's damn near magic and a great many here are kind of just taking it for granted.
Cool. Being sole proprietor of AGI 15 years ago should result in monuments and religions devoted to you today.
Cancer should be cured, and we should be a post-quantum interstellar fusion-powered civilization.
I wish the AGI crowd would finally shut up now that it's clear no one is even trying for AGI (OpenAI revised that to "$100B in profit")
What we're getting is incredible, where we're headed is incredible, but some people have such a fetish for futuretelling they can't just shut up and enjoy the ride.
This isn't the gotcha that you think it is: AGI's original definition is being able to do any task that requires human intelligence.
The unlock isn't AGI smart enough to invent quantum mechanics, it's suddenly being able scale human intelligence using grains of sand instead of decades of food and energy and nuturing.
All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Wait, what? Am I understanding that correctly? That sounds really bad
I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.
Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?
the bullshit machine is learning to optimize its bullshitting techniques!
<AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt
you cannot give this type of worker autonomy over anything.
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
"The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
I was actually wondering when they will release the new Opel Astra model.
Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.
On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
Brother they can't even release the announcement post cleanly without it constantly going down, they certainly wouldn't be able to release this new model without doing so in stages.
When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.
Material wealth is only a single type of wealth. Who's better off - the rich guy who's always yearning to be richer and never satisfied, or the lower income guy that mostly just cares about time with his family and is really happy where he's at?
They simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology".
It hasn't always been the case. Even then, having piles of money still does not gain access to the best military equipment. Sure, we've been living in a time where a couple people get to enjoy a wildly different lifestyle than the average, it just feels like it's about to be different in a way that isn't as ignore-able as someone enjoying a pina colada in a yacht somewhere
Mythos was never released. It's really just the writing on the wall. I'm not going to give up hope, but it's pretty hard to win a race when some people get a jump on the gun.
Being strongly on the AI saftey side of things what is happening was 100% predictable.
At first the race wouldn't even be noticeable. Then people would see things speeding up, for example hardware getting more expensive. Then when the capabilities really got useful most people suddenly realize the race is moving 1000 mph and they are never going to catch up.
> GPT‑6 Astra brings together years of research and big bets across pre-training
Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
It's more like OpenAI vs no one, at this point. Anthropic has shown they don't care about general consumers or small/med businesses. You can't even use their models without it giving refusals on the most mundane tasks.
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
Very well said. It kinda describes how unrealistic these expectations are.
Vibe coders want a model that makes them rich, without having any actual specific idea.
They write a very ambiguous prompt and expect to be amazed by the result.
The complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.
Lmao, come on dude, anyone whos used these tools for research knows it makes them lazier, less interested and dumber. You really want disease researchers become sloppers too?
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
From what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
To be, or not to be, that is the question:
Whether 'tis nobler in the mind to suffer
The slings and arrows of outrageous fortune,
Or to take arms against a sea of troubles
And by opposing end them. To die—to sleep,
No more; and by a sleep to say we end
The heart-ache and the thousand natural shocks
That flesh is heir to: 'tis a consummation
Devoutly to be wish'd.
...
And thus the native hue of resolution
Is sicklied o'er with the pale cast of thought,
And enterprises of great pith and moment
With this regard their currents turn awry
And lose the name of action.
Anthropic don't care, they don't want their products to be used by general audiences in any serious manner. Their interest is in selling to megacorps and using the models for themselves internally to swallow industry, and drumming up AI fear to regulatory capture to shut down the businesses that do want to make AI accessible to the people.
Yeah, but if OpenAI provides a cheaper model with the same capabilities the megacorps will buy from OpenAI. If you want to target enterprises you need to have some competitive advantage, it's not enough that "I wanna target them".
Yes. I had Codex rewrite and fix all of this in one shot earlier today (using Typescript). Unfortunately, I can not show you the code, because I do not know how this "git" program works but the AI keeps talking about it.
Is it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.
There will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do.
There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype.
There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.
But thank you for spending other peoples money to give us the tech regardless!
All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
I can't help thinking "doesn't matter much unless it's perfect" because if someone is using this to build a board (cool) but then it's not flawless, troubleshooting will be quite tough as a novice. Like, say, when I start digging into the web code generated by a coding agent.
I am most excited about it bringing down the barrier so more people join in on hardware fun, so hopefully it will unlock folks that stayed away in the past.
Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
That's a policy and distribution problem, not an AI problem. Anthropic is doing their best to make it a reality though. They'd love nothing more than to shut down distribution and become the sole gatekeeper of everything AI.
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273
How about we stick to that one for talking about the rollout, and this one for talking about the model?
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Take it from the mouth of the creator of ARC-AGI:
When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
That was 6 months ago, so the progress that Astra represents happened about 2x faster than I anticipated. I think the speed of progress will surprise a lot of people, and what the new models can do will challenge the views of AI that people developed by using prior generations of models.
I feel like AGI's definition got watered down, and these tests do not cover the original definition, what is your definition and thoughts on aligning with what all of us understood from the original claim?
I feel like this test is just helping someone like Sam Altman pretend like he implemented AGI as originally pitched for an IPO when in fact, he has not. Shameful.
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
- Sam Altman on AGI
I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine. Most of them, anyway. But I guess that’s “moving the goalposts”.
I can call coworker right now and have conversation so frustrating that I wish I was talking to machine instead.
Haha, sounds like a median human being alright!
Reminds me of Turing test. https://en.wikipedia.org/wiki/Turing_test
I feel like my odds are better with the AI than with random humans.
Random humans don't have a significant cross section of human knowledge available in real-time, although many like to pretend they do, especially in internet comments :P Being able to compete with the capabilities of a median human would be an absolutely world changing achievement.
> I can call a coworker right now and have a real time conversation with them without feeling like I’m talking to a frustrating machine.
These machines were designed with the express purpose of being condescendingly sycophant to the point they hallucinate just to state "you are absolutely right".
No wonder some people even find these chatbots to be wife material.
I dunno, have you tried the voice chat in paid ChatGPT?
Have you? It’s a dumb model that gets a lot wrong. Codex voice is IMO only good for when I cannot dictate into the good models. The delay introduced by letting it work with a good model kills it for me.
I think they're claiming it's achieved by text models, not voice models, fwiw.
Here's another definition of AGI from Sam Altman:
https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...
Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?
Kevin Roose (New York Times): I probably would, yeah. Would you?
Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.
So, "really good" is the boundary now, whatever it means.
If the new AGI benchmark is "be Einstein/Feynman" then we've hit AGI.
What if it can be Einstein, but can't draw a Pelican, write a solid college-level essay, or fold clothes?
The ability to do a ton of book learning in training, and pull in tons of related context at once, is superhuman in some ways, but lags a lot in others.
> What if it can be Einstein, but can’t draw a Pelican, write a solid college-level essay, or fold clothes?
Then it’s an expert system.
Stephen Hawking wasn’t very good at folding clothes.
The ‘General’ part of the term ‘AGI’ seems like a trap to me, because there will always be new workflows to master. Can Astra one-shot level completion on some yet-to-be-released video game? If no, does that mean it’s not yet ‘Generally’ intelligent?
You won’t get pure ‘general’ intelligence until you find Einstein’s hidden variables and load the state of the entire universe into context.
Meanwhile, building a series of expert systems targeting specific valuable workflows is useful today and seems like it’ll continue to scale to cover huge swathes of economically valuable workflows.
I think that’s the more interesting thing to be measuring. The surface area of useful economic workflows that can be addressed with expert systems built with today’s tech.
Hitting some ‘Artificial Expert Intelligence’ coverage threshold on economically valuable workflows is what will matter for humans well before pure ‘general’ intelligence.
The only important part of 'general' is the ability to learn from experiential data and update your own model. That's what leads to general capability. Humans can't oneshot any task natively, but we can practice for a while until we uncover often novel methods of accomplishing something.
Therefore: the current transformer architecture is fundamentally incapable of AGI because the models have no mutable long-term memory.
You only have weights (large immutable memory), or context (small mutable memory).
Humans have mutable long-term memory: I can learn a new skill, adapt an old skill to new information, or learn new knowledge today that I couldn't perform/didn't know yesterday. I don't have a training cutoff.
Context engineering is an attempt to paper over this limitation. You can get really far with context engineering and huge models, but you will never get to AGI because there are many tasks where humans' mutable long-term memory outperforms.
For example, a human can invent a new musical instrument and then learn how to play the instrument they just invented. That's inference (inventing an instrument) leading to training (neuroplasticity). Humans have the ability to train our NNs with considerably fewer training samples. Everything that you can do with transformers is in one causal direction: training -> inference.
Folding clothes is happening. https://www.youtube.com/watch?v=cRZNwgvcWUg
AI in math is ongoing. https://spectrum.ieee.org/ai-in-mathematics
Adding sibling comments, I think some people may be overestimating how well the median human can draw a pelican, or create an SVG of a pelican (depending if we’re comparing to an image generation model, or SVG generation).
Most people can't draw a bicycle. There was an artist 10 years ago that asked people to sketch a bike, and then turned these sketches into 3D renders - quite funny.
https://mymodernmet.com/gianluca-gimini-velocipedia-bicycles...
https://qz.com/681345/an-artists-3d-renderings-of-bicycles-d...
How good was Einstein at drawing pelicans on bicycles by writing SVG code?
Checkmate, meatbags.
Laundry folding has become a doable demo for startups, and ChatGPT has been spitting out college essays for years.
Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this: https://crimson-jeri-74.tiiny.site/
And the only reason LLMs can't write essays indistinguishable from human output is because they aren't RLHF'ed to write like humans.
Folding clothes isn't an LLM's job but if you were to insist, they could certainly do it, as any number of videos from robotics labs will attest. That particular future is already here but definitely not evenly-distributed.
> Pelicans are a solved problem at this point. An open-weight model on my own machine gave me this
That feels kinda like when I remember seeing Ocarina of Time for the first time, and thinking “oh my god, this looks just like real life…”.
For me it was the wheels. I couldn't stop staring at the wheels... how did it get them so freaking perfect? Mad respect to GLM 5.3.
I can't draw a pelican. Literally my only point of reference would be AI pelican drawings from the test. Otherwise I wouldn't know how to draw one at all.
I would be able to draw an accurate bicycle, but I'm an outlier on that. Most people could not draw one [1].
[1]: https://www.booooooom.com/2016/05/09/bicycles-built-based-on...
Wouldn’t that mean producing novel work like relativity and QED?
I would maybe argue that Einstein was the most LLM-like of great thinkers.
A lot of his great discoveries were mostly that he was very knowledgeable about the bleeding edge research in a number of disparate areas, and was able to have the aha moment where he could make the connections for how to integrate them.
A lot of other thinkers who created new fields from scratch are probably way harder for an LLM to crack.
That is very aligned with an LLMs ability to have superhuman knowledge in wide areas.
What is novel physics?
I think they mean improve our understanding of physics with new theoretical results or paradigms. Like if it’s 1899, would Astra develop General and Special relativity on its own?
This is as good a time as any to note that we might be closing in on a new conceptual revolution in our own time as it relates to holography and an information centric approach to spacetime. Obviously It's the furthest possible thing from a guarantee, but it has much of the enthusiasm and motivation that string theory had previously enjoyed in prior decades.
So it could be a natural experiment for whether AI can contribute to novel physics. Specifically, there's a big question about weather. Something like our informational understanding of black holes where information inside it is equivalent to information on its boundary (which I'm sure I'm not saying correctly), might be generalized to regular space-time. More people should be freaking out with excitement about this and perhaps it's something to which AI can contribute.
Do you have anywhere you recommend where I can read more on this?
Honestly I don't think I have single great article, though some Quanta ones are ok, and the Wikipedia article is okay.
The best thing I can recommend is what I did, which is ask Claude about the significance of (1) quantum computing error correction, and (2) error correction in black hole holography and research convergence between the two.
https://en.wikipedia.org/wiki/Holographic_principle
https://www.quantamagazine.org/how-space-and-time-could-be-a...
Edit: this whole article, despite it's boring title and hook, is maybe the best discussion of holography as a recent and active research frontier.
https://www.quantamagazine.org/if-the-universe-is-a-hologram...
There was symbolic AI programs in the 1980’s that “discovered” Kepler’s laws and the resulting solar system model from just tycho brache’s astronomical observations. That was the the very first “new physics” ever.
i wonder if we could train a modal, and omit all data prior to 1899, and see what happens?
Do you mean after? People do this!! But I think it’s a bit different. It won’t be apples to apples because the data volume I think is just so much different. Maybe there are good experiments for something like this.
as always, as good as its prompt...
I assume solving one of the major open problems of physics?
Would this be possible without it being able to run novel real-world physics experiments autonomously?
(Note: I am not suggesting we let it do this. Please don't, in fact)
AI could discover candidate novel physics without autonomously operating new physical experiments, and humans or instruments can later independently validate the result. This is analogous to how Einstein developed theories whose predictions were confirmed by experiments and observations only years or decades later.
Do we have any examples of an current day AI system introducing a novel concept or perspective. We've got plenty of counterexamples discovered and some theorems proven, but afaik nothing analogous to a new definition.
Pharmaceuticals already are…
It could be a good theoretical physicist. Actually it could be a good experimental physicist as well since senior experimental physicists use grad students for the manual labor.
Why not?
That's how we get AM.
Mornings? Long-wave radio?
https://en.wikipedia.org/wiki/I_Have_No_Mouth,_and_I_Must_Sc...
That’s already been done. I know of at least one novel result contributed by Claude to frontier physics. I’m sure there is more.
For it to be like a human it wouldn't just need to solve existing phsyics problems, it would need to push the field forward and introduce new paradigms.
Solving "open problems" will push the field forward.
My comment wasn't very long, yet you somehow still ignored the main part, "and introduce new paradigms". The point is whether it can do everything humans can, entirely new theoretical frameworks and ideas, such as string theory or dark matter, are not coming out of AI at the moment.
Creating new physics is the new AGI goal post
If it makes you feel any better the curmudgeon who drives the ARC-AGI tests feels the same way, that's why we're on 3 and I'm sure we'll see 4. Also, we can all avoid calling it "moving the goalposts" so no one feels talked down to.
It’s better to call a spade a spade.
> I feel like AGI's definition got watered down
Typical result of venture capital and too many bag holders unfortunately.
People need to stop redefining and trying to capture the term AGI. None of this is AGI. Not even close. Can it do things an AGI could do? Yeah some of it, but the difference really matters. These things still regularly fail, and gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
An AGI wouldn't struggle with that.
The last version to fail on those questions was GPT 4.5.
Meanwhile most humans fail to correctly answer how many f's are in the sentence, "Finished files are the result of years of scientific study combined with the experience of many years.".
It comes and goes... My point is we're not near AGI.
Nooo. I just asked Claude (Sonnet 5 Medium) how many I's are in assassin, and it said 2. Granted, it got several words correct before this. But no, they still aren't great at this letter-counting thing.
I don't think we should count the lower tier models if we're discussing what the top ones are capable of. No one was suggesting that Sonnet is AGI.
Meanwhile Qwen3.8 27B got both the 'f's and the 'i's questions right.
Anyone know why they aren't good at this?
Tokens are the most basic input unit of an LLM. But tokens don't generally correspond to words or letters, rather sub-word sequences. So Strawberry might be broken up into two tokens 'straw' and 'berry'. It has trouble distinguishing features that are "sub-token" like specific letter sequences because it doesn't see letter sequences but just the token as a single atomic unit. 'Straw' and 'r' are two tokens but an LLM is entirely blind to the fact that 'straw' has one 'r' in it.
As an analogy, I might ask you to identify the relative activations of each of the three cone types on your retina as I present some solid color image to your eyes. But of course you can't do this, you simply do not have cognitive access to that information. Individual color experiences are your basic vision tokens.
LLMs see tokens, not words spelled out with letters.
Imagine verbally asking someone who has never seen written text the same question: unless they memorized the answer for the specific word you're asking about, they'd have to guess.
People assume that's the reason because it's intuitive and "strawberry" is one token. But that doesn't explain why those models would also often get it wrong for "StRaWbErRy" or even "s-t-r-a-w-b-e-r-r-y", where the r's are not combined into one token.
We don't know what was going on inside the closed source GPT models, but this paper investigated on some of the open-weight models and found it's not due to tokenization: https://arxiv.org/abs/2604.00778
A lot of movies gave the impression that making phone calls and balancing checkbooks would be the easiest tasks for consumer AI to solve, while math and science might require exceedingly advanced AI. Turns out to be the opposite: computers are great at math and bad at conversation.
That's because movies were based on the "general" nature of AI, assuming we would create intelligence that would learn and grow.
Pretty much the opposite of what we can't up with, if you're willing to call what we have intelligence.
I'm going to go out on a limb and guess that there are plenty of savants who can't tell you how many r's are in strawberry.
These are like saying someone isn’t human because they have a speech impediment or an auditory processing disorder.
AGI doesn’t mean infallible, it just means it can have a reasonable crack at things it hasn’t seen or done before.
I don't understand why you're being downvoted... that's literally the definition of AGI.
> gaslight about answers to questions like "how many r's in strawberry" or "s's in espresso".
this has been debunked too many times to bother rebutting. they struggle with those things because of the way they are.
it's completely irrelevant.
If it allows you to distinguish readily between human intelligence and computer intelligence, it seems quite relevant to the question of whether computers have achieved something akin to human intelligence.
It may not be useful for anything else, but at least it can say that.
But it's not relevant as a metric to gauge distance to human intelligence. Humans see individual letters, LLMs do not. If I asked you the relative activation of the cones in your retina as I showed you some solid color image, you couldn't do it. You simply do not have cognitive access to that information. But that says nothing about your intelligence.
A more accurate test would be to give it a list of words (or anything represented as a single token) and ask it how many times that token appeared. I'm sure they have no trouble at that task.
but a human doesn't attempt to make up an answer, the human knows that he doesn't know?
Yes, they have some alien failure modes. But that should be expect, they are an alien intelligence. I might be willing to grant that a lack of ability to reflect on its own level of knowledge is a demerit to it being generally intelligent. But then again it is largely an artifact of training. I suspect if there were a guessing penalty during pretraining they would develop or more readily communicate the strength/reliability of their knowledge.
this is so false that Dunning and Kruger invented a name for it
I am not enthusiastic about criteria for human-like intelligence that imply that dyslexic people don't have human-like intelligence.
[EDITED to add:] I actually don't know whether dyslexic people find it difficult to count letters in words, if they have them already written down by someone else. I suspect they find it harder than people who aren't dyslexic. But perhaps "blind people whose spelling is poor" would have been better; I would not want to deny them human-like intelligence either.
I have slight dyslexia. I can't automatically write double consonants all the time.
But counting certain type of letters is very simple algorithmic task especially in written text. Any reasonably intelligent actually thinking thing should come up with algo and then execute it. Which to me sounds like reasonable minimum bar for general intelligence.
which just brings us back to the whole birds vs planes thing.
turns out that flapping wings is not the right way to unlock human flight.
computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.
> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.
I don't think it's irrelevant but perhaps not in the way you're assuming. When assessing AGI I'm not evaluating counting characters or even the execution of math operators at any scale or speed. As you observe, computer software from Regex to spreadsheets and Mathematica already handle that well. But AGI isn't about what computers can do, it's about whether AIs can do the specific things which, until now, have been uniquely human capabilities. Like understanding nuanced context and then coming up with novel approaches to solve a new kind of problem not relying on any specific prior training or knowledge (the 'G' is for General).
Most definitions of AGI start from a baseline that already assumes easily passing a Turing test and doing anything via text response that a high school graduate could. I ding LLMs not for failing to count but for failing to intuitively understand the nuanced context of a simple class of problem it hasn't seen in its training data. I fully understand that the reason LLMs fail letter counting is that they operate at the token level. They weren't trained on individual letters first, like human 2nd graders.
The only reason recent LLMs get strawberry and blueberry correct now is that they have those words on their pre-training 'cheat sheet'. However, the underlying fundamental weakness in the way LLM intelligence works which leads to this failure mode still hasn't been addressed. Even when the frontier labs add "recognize any sub-token counting question and write a Python script" to the training cheat sheet so LLMs always pass that test... they'll still be unable to recognize a simple class of problem which isn't on their 'cheat sheet'. As long as that's the case, to me, they aren't AGI because they can't fully replicate human-like recognition of novel problem classes. And it's not just about letter-counting. That gap and others like it lead to many other kinds of non-human brittleness in LLM problem solving. Those are the classes of reasoning, intuition and insight that the ARC-AGI series has been trying to queue up as targets. Not to show how bad LLMs are but to help them be great in all these counter-intuitive edge cases
Models often write python scripts for counting such things…
appreciate your response, but it's still birds vs planes.
AI does not need to feel emotions or have a heartbeat to be useful. It only needs to perform a task correctly à la Chinese room.
>therefore cannot fully replicate human-like intelligence
this does not follow. planes don't flap wings therefore they cannot fly?
> planes don't flap wings therefore they cannot fly?
This example still misses my point, which isn't related to usefulness or economic value. I concede that LLMs can have greater utility and economic value than humans on many tasks. The point is most definitions of AGI include something like "can fully replicate all the routine daily tasks done by any competent high-school graduate." That's not related to whether LLMs can solve many high-value problems faster and at larger scale than any human. That was also true of ENIAC in 1946.
The fact an airplane can fly faster and farther than any bird is irrelevant to whether an airplane can "fully replicate all the routine daily tasks done by any competent bird." That's the bird equivalent to most AGI definitions. An airplane can't build a nest or recognize the signals encoded in birdsong.
In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any". And in this context, airplanes scoring 15,000% more than birds on 'speed' and 'distance' doesn't matter any more than AIs scoring 15,000% more than humans on 'add 10,000 numbers'. We still aren't near AGI because LLMs cannot fully match any high-schooler's ability to independently conceive new approaches to novel problems not in their prior training data.
> isn't related to usefulness or economic value
which gets us closer to philosophical questions which I'm personally not that interested in.
>In the same way airplanes fail the 'bird replacement' requirement, AIs currently fail most AGI requirements only on the terms: "fully", "all" and "any".
I'm not sure we want a machine that fully succeeds that test.
Planes pass the 'bird replacement' test on the only criteria that matters to us ... flying.
If we wanted nest making planes I think we'd have them by now. Nest making doesn't rate highly on the problems we're looking to solve though.
I don't want a machine that is moody, or depressed or has schizophrenia, which are all pat of the human condition.
We don't need the human "intuition magic dust" to do 99.99999% of useful work.
They're machines designed to do the work we don't want to. That's as "general" as their intelligence needs to be.
I'd prefer if my clothes folding machine did not have an existential crisis.
> which just brings us back to the whole birds vs planes thing.
That just says we don't need to design an AI like a brain. That's not part of this discussion at all.
> computers could count the Rs in strawberry since vacuum tubes. that measure is irrelevant.
I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"?
The fact that very basic computers can do it makes failures embarrassing when testing for AGI, not irrelevant.
>I'm confused, is your argument something like "It's too easy so AGI doesn't need to be able to do it"?
Do you possess magnetoreception? a stupid pigeon can "see" the earth's magentic field. why are you blind to it? does a lack of magnetoreception make your intelligence any less "general"
no, you're just blind to it because that's just the way it is.
LLMs are blind to character counting because that's the way they are.
It didn't stop ChatGPT from finding the Jacobian Conjecture counterexample.
Human intelligence and machine intelligence are only going to cross over to a certain degree.
same as plane flight and bird flight are only kinda related.
> LLMs are blind to character counting because that's the way they are.
But if I can't calculate it myself I know to use that basic computer to do it, not make up an answer.
> Human intelligence and machine intelligence are only going to cross over to a certain degree.
That's where the word "General" kicks in. If there's big limitations on the overlap forever, then there will never be AGI.
>If there's big limitations on the overlap forever, then there will never be AGI.
maybe. we'll see.
It matters under the lens of AGI. Artificial intelligence that can meet or exceed human intelligence. Sure, it's good at some things but rather limited at others.
Breadth of capabilities matters... and a promotional video is nice and all, but people are throwing this term around like it's a prize they've won, but they've not gotten there yet.
> they struggle with those things because of the way they are. it's completely irrelevant.
I mean, they seem like fair game if you’re ever participating in a Turing Test.
I wonder if Altman's definition also includes taking on the same liability as a coworker would.
Probably not.
Would 100% need to be fully insured for liability, with a sizable war chest that OpenAI cannot even afford.
This implies that human employees don’t have insurance. But they do. My company’s cyber insurance for example covers breaches due to employee mistakes. Most companies also have umbrella liability policies. It’s just that for now, AI “employees” need a LOT more coverage.
And that's the rub, isn't it?
If it can replace a worker but does too much work to be checked routinely by a human, and bears no real responsibility for its actions, well... it's really just a way to jack up the value of the settlement the company using it gets to pay out when it does something that causes a lawsuit.
If OpenAI had just simply stuck to making "good enough" models that were open sourced (like they promised they would be when starting out) and could be used to augment a human doing a task - a human that could be given actual consequences for messing up - they wouldn't have burned all of this money trying to reach this nebulous definition of AGI. Hell, "good enough" is what many open-source models are, and that's what terrifies Altman.
What do you mean by bearing no real responsibility for its actions? If I use a model to accomplish a task and it fails, I use something else to try to accomplish the task. If it's my responsibility to complete the task, it can't be the model's responsibility unless I've agreed to some sort of guarantee from the provider.
If the provider says "the model will always be right or your money back" then the provider has got responsibility. If there's no guarantee, there's no responsibility on their part, just on the person whose job it is to try and solve a problem with the model.
The "dream" that these labs are mostly selling is the ability for capital to subscribe to their AI for cheaper than it costs a human to do some task. Not to have a "human + AI hybrid where the human is responsible". It's what the whole AGI valuation is based off of, in that scenario, with no human oversight, the agent they lease has to be responsible for the task?
It's not really a dream though? You can subscribe to their AI right now and complete many tasks for cheaper than it would cost to pay a human to do those tasks. Another person can do the same thing but also keep a human in the loop. You and the other person may compete in the market for whatever your product or service is, and you both might do very well or one of you might do better than the other because of a whole host of different reasons. Nothing in the process of developing and selling access to a more advanced LLM requires the customers to do away with human labor, nor does it require them to offer the LLM service with a guarantee that it will never make any mistakes. So far the LLMs have always made lots of mistakes and the companies sure keep making a lot of money.
The person responsible at that point is the sucker who fell for the dream.
> What do you mean by bearing no real responsibility for its actions?
If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.
It it makes a mistake and deletes your website from AWS, who is responsible?
If it targets another website because it decides that it is "part" of your website and attempts to break into it, who is responsible?
> If you give an "intelligent agent" offered by one of these model providers a task of updating the content of your website, and it updates it with inappropriate adult content, who incurs the cost of the machine's error? The model provider generally does not.
Something in your prompt led it to do that, is alex0015's point. The statistical odds of these frontier models screwing up to that extent are so impossibly low that it would almost have to be intentional or accidental negligence on the part of the prompt writer to accidentally have their agent write pornography to their website.
The burden of the mistake would have to fall on the person that gave the tool instructions, because it can't know that what it did was wrong. Wrong is subjective in this case. It only did what it did because you, figuratively speaking, encouraged it to.
Yes but I'd actually go even further than that. We've all had experiences where the model straight-up produces gibberish sometimes, right? It's happened to me with a badly configured harness on a local model, and even also on frontier models like when you used to ask them something innocuous about the seahorse emoji.
What I'm arguing is that yes, it's your fault if you misprompt the model and it does something catastrophic. It's also your fault if you prompt it correctly and it does something catastrophic anyway due to some other glitch beyond your control.
The whole discussion started with the point that OpenAI could be financially responsible for damages if their models cause problems for users. If my job is to design a system that works, and instead my system doesn't work, it's my fault regardless of whether I used no LLM, a local LLM, or OpenAI's LLM. Depending on whoever's in charge of doling out consequences, I might get away from it with zero, light, or heavy consequences. But at all levels, I can't reasonably expect to deflect blame onto the model itself.
In all of these cases, it's you. It would be the same if you downloaded an open model, ran it locally, and it happened to make the same catastrophic mistakes. The consequence to the provider is that if they offer a product that does these things, people don't buy the product.
In general, the person whose job it is to provide the company with a working, non-adult website and not hack into other websites is the one who would receive consequences for failing to meet those expectations.
The problem I see here is that ultimately, you'll have capital wanting to replace workers like others have said, and have someone roughly equivalent to a manager or vice president driving teams of agents to achieve business outcomes.
These tools can push out more results than a human can hope to evaluate in a business-sensitive, or even realistic, amount of time. You have to take it at its word that it did things right, and there's no real fear of failure or consequence on the behalf of the agent.
Of course, but "capital" is no stranger to risk management. I'm sure we'll see some spectacular failures, but most will handle this just fine.
If future jobs are simply reduced to liability scape goats (or more appropriately reverse centaurs) for management to pin things on then I'm taking up goose farming.
I'm afraid you will be pushed out of the goose farming market by these new ultra-efficient farming bots.
That's more-or-less what you are now, especially if you work at a company like Meta where 1) the guy at the top holds majority control of the company's shares and 2) keeps making massive, expensive mistakes either by accident or design.
ARC-AGI was never "if this benchmark is saturated, we're at AGI". It was always about crafting adversarial tests that humans are good at, but current AIs are bad at. Point out the gap, get AI teams to attack them.
In practical terms? They usually get solved with a bigger badder LLM. "New ideas are needed?" Nah - ten times the params, ten times the test time compute.
ARC-AGI-3 was more of a failure in that regard than -1 or -2, because even on day 0, an off the shelf LLM with a harness could get 50%+. And messing with evals by forbidding "LLM with a harness" from scoring? Yeah no, that was just bad.
>When we released ARC 3, I got asked, "when do you think a frontier model will saturate it?", and I answered "in about a year, though it depends on how much it gets explicitly targeted"
You're treating an off-hand comment by an ARC 3 researcher as some sort of a precise AI capability acceleration benchmark. Can we leave casual anecdotes (even from researchers) out of the discussions please?
François Chollet wrote in February that he expected ARC-3 to be saturated in "about one year".
"Frontier models today perform very poorly with a minimal harness. However if big labs start directly targeting the benchmark like they did for ARC-2, numbers will go up fast."
https://x.com/fchollet/status/2022054537293705260
Well it's not exactly saturated when OAI refused to use the harness explicitly provided by ARC-AGI. I'm not really familiar enough with the benchmark to declare whether it's a perfect measure for AGI, but I kind of doubt it is.
2x faster at what resolution? is 6mo vs 1 year really that different? Usually surprise comes in order of magnitude mismatches in expectations.
Yes, 6 months vs 1 year is huge for technology that has gained wider adoption only recently.
Adoption means nothing.
2x gains from a mature technology would be surprising.
2x gains from a new tech would still be called “low hanging fruit” in another setting.
I don’t read enough to know in what ways the training / other technical steps have really advanced.
You'll also need to compare the amount of compute used now and then, which seems exponential to me.
Are there plans for ARC 4?
We are nowhere near AGI. They all talk the same, they can’t help but try to please and affirm us, and if you engage them for too long they become incoherent. They are facsimile machines. They are xeroxing language - but not even, because we can’t even duplicate our results. Too many people mistake the black box quality for magic.
Super useful, incredible tools, but not AGI. Try and roleplay a dialogue with one, make it whatever character and scenario, and see if it can sustain a coherent conversation for more than 30min with you AIM style (aol instant messenger, if that isn’t clear). Expert mode: never correct or adjust it mid conversation.
I’m not even talking about repetition and predictability. It’s nothing like talking to a person. And in a short amount of time it literally can’t form a coherent sentence.
Human brains have difficulty reasoning about exponential growth.
They keep saying that. I'd say it is more like human brains that don't remember high school math have trouble with it.
If something at rest is accelerating at 9.8 m/s^2, how long in seconds will it take to reach 10% of c? Answer to the nearest order of magnitude - will it take approximately 1000, 10k, 100k, 1000k seconds?
I’m sure you know this is an exponential growth question but have no intuition of the answer.
That is a linear growth problem whose answer is very easy to intuit.
You must be joking. A high schooler with a few hours of physics classes can intuit the answer.
Until special relativity kicks in to completely invalidate whatever intuition you have about this problem.
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
To me AGI is all about the "G" general (we already had the AI part). General meaning universal, everything. It's not a function of knowledge or specific hardcoded tests, it's that you could give it a test it's never heard of before and never been trained on and it would ace it (it might need a lot of time).
Currently LLMs can't even really learn within a conversation, they can add a note to context and try to not drop it. Example things an AI cannot do yet (but maybe someday will):
- write a well-received book, write a best-seller
- come up with a new company idea, Run that company
- actually have a decent conversation, maybe someday talk somebody out of suicide effectively
- come up with its own ideas or theories that nobody else has presented
- understand the stock market well enough to trade better than an index fund
- be an expert Game Master in a TTRPG (making no mistakes, getting a read on the players' fantasies, calibrating difficulty in response to emotions)
- come up with a theory of what makes games fun, make a popular game
- be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
- be able to articulate what it knows, what it doesn't know, and what information it would need to have to answer complex queries
- exhibit metacognition (thinking about its own thinking) and self-optimization
- wonder about things
- observe contradictions and ironies in the social-consciousness, do a standup routine that makes you rethink how you look at things
To me all this makes the label of AGI completely meaningless.
What AGI has always meant (eg. in 2019) is Artifical General Intelligence.
Artificial -- something made by humans instead of occurring naturally
General -- not confined by specialization or careful limitation
Intelligence -- the capacity to learn, reason, solve problems, think abstractly, and adapt to new situations
Basically, the metric was that any healthy adult human on the planet represents a general intelligence. This has certainly long been reached.
Also some of the stuff you're listing has long been solved as well, such as listing what it knows and what it doesn't know, and what information it would need. Other is just poorly defined: "be able to argue persuasively". AI can certainly write an argument on almost any topic that would pass any University homework in 2019.
By this definition, even most humans would not qualify as having AGI though.
Humans are AGI simply because if they want to, they can achieve it (it just takes some effort).
However most humans can do at least some of the things given they spend the required effort.
Some problems presented needs a very large context and some are not much solvable (e.g. trading) since market responds to traders' actions, as well, making it effectively an oracle problem (of computation).
On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them. However, brains in nature never stops. Wonder, daydream, sleep, self-evolve, clean up and eliminate memories and views and much more.
But this is assuming the model is the entire story. The original comment you were replying to pointed out that the harness is just as important.
The hardware of human intelligence is not a singular thing that is uniform throughout. You cannot take the prefrontal cortex white matter out of someone's head and say you are holding a person. Much of the parts of our brains that enable much of our intelligence, is made of different specialized stuff. The visual cortex and sensorimotor regions aren't only there for input and output, they are used by the more thinky parts of the brain to do visualization and spatial reasoning. The cerebellum contains billions of neurons making little oscillator circuits and PID-like self-regulation machines that help make muscles do what they're supposed to, but also provide attention and time perception.
Heck, our brains contain language models, that train themselves up based on a glut of data over a span of about 10 years, and then they become more or less set in stone for the rest of our lives. Of course we can learn languages, but the "Critical Period" is a very real thing that produces a permanent architecture for some grammatical structures, or things like the ability to partition a lexicon by gender for faster lexical access which cannot be learned as an adult if your native language did not have gender.
I'm not trying to make a direct analogy, the point is that the language model doesn't need to be fully "generally intelligent" all on its own for there to exist a general intelligence, because the language model can be part of a generally intelligent system, which can do things like form, recall, and manage memories which are by now a standard feature in basically every chatbot.
> On the other hand, we must be aware that these models are static, and they indeed stop when nobody asks something or requests an action from them
The parent commenter noted:
"if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI"
Harnesses absolutely can enable models to continue thinking about things. And LLMs do wonder and explore weird ideas like daydreams when you allow them to do this.
I'm guessing their defn of AGI is something like the sum total of all humans' abilities? Still though some of those tasks (e.g. beat an index fund) may very well be impossible, and worse yet a lot of those tasks are not coherently defined.
I haven't met a person who doesn't wonder about things.
They can't, by definition, have the A part btw.
AGI has always been expected to outperform humans or else what is the point of it?
- Amazon is full of AI books, and they're clearly making money. AI has won multiple literary and artist awards.
- Okay, it's not a "new company" idea, but VendingBench is all about ability to run a company
- Plenty of people disagree with you on conversational quality; see "AI Boyfriends" etc.. (and it's not hard to find people who consider it uniquely valuable for discussing mental health)
- "come up with its own ideas or theories that nobody else has presented" C'mon, seriously? Solving a half-dozen hard open math problems wasn't enough there? What the heck counts as "it's own ideas or theories" at this point?
- plenty of evidence that custom models are starting to do well on the stock market, although I'll admit we're a year or so from any solid proof, since you need a track record to really make the claim
- LLMs have been capable of being a GM for a TTRPG for over a year (although like humans, they make mistakes)
- Okay, conceded, but humans tend to take years and large teams to make a game. Even if the capability existed today, it would take a while to actually build, test, market, etc.. - all made much more complicated by gamers being largely opposed to AI art styles, etc..
- "be able to sort through research and come to conclusions on complex geopolitical/sociological topics" - uh... did you mean to say something else, because "come to conclusions" is... like, LLM 101?
- Hahaha, have you met humans? We definitely cannot do that.
- Uh... thinking about it's own thinking is trivial. Most LLMs these days are built using LLMs, so uh, self-optimization seems nailed, too? We just don't let them do it unsupervised.
- LLMs fucking love to wonder about things
- "observe contradictions and ironies in the social-consciousness" really seriously have you actually used an LLM recently? I think you would find it remarkably enlightening.
It also cannot do tasks it wasn't trained for. It can extend texts, read images and click on a desktop, but only because it's made for that.
I don't think that's strictly true, as I can give it a new gui or tui program it wasn't trained on and it will learn it. Unless you're talking about general abilities like sight, but the same is somewhat true of humans.
If you consider the data on which an LLM was trained on to be points on a very highly multidimensional object, the claim is that the LLM can interpolate a convex hull spanned by those points, therefore recovering a subset of consequences attainable from those points. Obviously this hull includes completely novel points that were not present in the initial data set, so the output of the LLM goes beyond its initial training. And yet, there are clearly points outside a convex hull spanned by any finite number of points, such that we can imagine not all possible outputs are attainable using this method.
The claim is furthermore that truly original thinking, the infamous leaps in understanding and creativity, happen by attaining points outside such a convex hull.
It's hard to rigorously verify or disprove this claim. Hopefully this helps build an intuition of why the claim is not as shallow and obviously wrong as it may seem initially.
I'd call it interpolation on a high dimensional manifold. Convex hull is too simple a shape.
But yes, metaphorically I think that's right.
This makes me realize there is a higher bar we need to achieve with AI still. The ability for the model to evolve through interactions more on a hourly or daily basis. The models are accelerating but inference doesn’t modify the model.
Humans cannot do tasks they are not "trained" for.
speak for yourself
>be able to sort through research and come to conclusions on complex geopolitical/sociological topics (e.g. theorize on whether AGI will result in mass poverty or mass abundance and be able to argue persuasively)
i'm not sure what makes you think AI cannot do this already. in my experience, this sort of deep research is something AI is quite good at.
example i just tested: https://chatgpt.com/share/6a9a20e3-1d20-83ea-a125-31aa240c74...
I understand that what it came up with sounds impressive (especially since I know 0 about Myanmar), but on the topics I do know about its analysis routinely have very fundamental problems (even this Myanmar analysis has % that add up to > 100). There's a chance it's just parroting the majority opinion on Myanmar, or making stuff up (and perhaps you could ask it to write a strongly worded opinion in the other direction that would sound equally plausible).
For example I asked it to do a full analysis on the AI bubble, and a full analysis on the risks of Glyphosate, and it came up with a lot of things that sounded credible, but within a few minutes of questing was admitting it hadn't even really checked for internal consistency in its positions, and even doing a 180. It certainly was much faster at gathering sources and reading but it fundamentally doesn't seem very effective at creating a consistent worldview.
And of course the funny thing is it says it did a 180 on one of these topics, great, except whatever it concluded will be discarded because it cannot learn. It's just bonkers to me pretend this is AGI, it probably couldn't even hold its own in this very discussion.
I claim that the percentages add up to more than 100% because the first described case overlaps with the second.
Your examples are things that most humans cannot do, or things that AI can already do. For example most humans, even most intelligent humans, could not write a well-received book, run a successful company, or make a popular game. On the other hand, AI can absolutely sort through research, draw conclusions on complex topics, and argue them persuasively. Likewise, I don't know what you mean by a "decent" conversation, but millions of people converse with chatbots daily, so I don't know why you say AI fails to meet that bar.
I fail to see how what you describe is any better than the old autocomplete-on-steroids comparison. Could the human mind learn to spell every word properly? Yes. Do most (or any) do it? No. Does that mean a human can't?
I think what OP was drawing a comparison to is that AI right now could not come up with an award winning novel from the spark of some creative notion and working up from there, as opposed to just mashing together what has already been done and calling it a day.
> AI right now could not come up with an award winning novel from the spark of some creative notion
I agree, but in this field we value evidence. So there needs to be some test of novel-writing abilities.
Once there is, AI companies will be out to score highly on it.
Wait for a resurgence of Philip K Dick-style novels as humans desperately try to write things LLMs cannot.
> Wait for a resurgence of Philip K Dick-style novels as humans desperately try to write things LLMs cannot.
I myself can't wait for Finnegan's Wake 2
Most of humans don’t reach any of these levels.
But you have to acknowledge how uneven the playing field is. The AI has read every book that's ever been written, and can spend hours of compute time in a few seconds.
I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund.
What I'm pointing out here is that these models appear to be intelligent when they really are simply unimagineably knowledgeable. When you drop the time-constraints it starts to become more and more apparent that human intelligence scales better with time than AI does (much in the same way AI can burp out tons of code but make your codebase entirely illegible within a matter of months).
Perhaps to simplify: my notion of intelligence is how much can you deduce with a constant set of starting context
> I think if a person had those same advantages (e.g. could spend 5 hours thinking about what to say next) we could all hold outstanding conversations, or if we had read every book ever written I think many of us could write a very popular book, if we could read every singe company's P&L statement in a few seconds we could invest better than an index fund
I’m not so sure of that - to get average outcomes in these fields it’s a matter of time, to get above average or extraordinary, you need talent/intelligence/taste.
And the bar the parent set is at extraordinary.
I would be willing to bet that any human for which we spend $100billion - $3 trillion (depending if you want to count single corporations or global totals) on in an attempt to make them as capable as possible would be able to reach all of those levels.
I think I view humanity fairly positively, but I admit I would gladly take the other side of that bet
A quick search suggests that the most expensive education in the world is something like $100k.
So like you spend a million times more than that and you still think you're not going to see some results?
> A quick search suggests that the most expensive education in the world is something like $100k.
I have a kid in an American university right now, and a quick search of my bank account statements confirms that there are far more expensive educations in the world.
Wow, an increasing number of US universities are going over $100k per year. That's a crazy amount. That's more than enough to hire an entire person.
Some results, sure.
But that list is extremely ambitious. Write a best seller, make a popular game, come up with a truly novel theory, consistently outtrade index funds.
That's top 0.001% human stuff, I don't think you can take just any person and get there through education alone, it takes extreme talent and dedication. There's also diminishing returns when spending on education, it doesn't just improve linearly.
The marginal return on education spending decreases fairly quickly, but obviously becomes zero at the point by which there are not enough hours in the day/year/decade to cover every single topic that humans know about - no matter the talent or resources available to the student.
except you are missing the one versus many argument here. sure we could make one human much smarter, could we make endless copies with the same intelligence? no
For $100 billion we could pay ivy-league level tuition for a million people. You don’t think investing that much in education would yield some good research or companies?
That's not what the parent comment said though
> any human for which we spend $100billion - $3 trillion...would be able to reach all of those levels
Are you imagining artificial augmentation somehow? Purely through tutors or training programs we seem pretty limited. Otherwise billionaires (or even multimillionaires) could have far more consistently successful kids.
Don't the children of the wealthy famously have a tendency to be successful? Or have I badly misinterpreted the last several thousand years of human history.
To really drill down into that I would think you would need to figure out how many millionair children get tutored vs how many get spoiled.
Successful, sure.
Gold medal Olympic athletes who are also brain surgeons AND astronauts, no.
You'd get rapidly diminishing to zero returns after the cost of university a few times over. Every dollar past that would produce no performance gain beyond that.
You want a computer program to be able to take a single phrase and execute decade long journies?
Who will be responsible for the outputs and side effects of such a closed loop system?
Half of those the agent fleet systems can do right now.
These are things it cant do and will not be able to do without human labor and long running human vision:
https://rcsnyder.github.io/open-frontier-curriculum/05-front...
https://rcsnyder.github.io/open-frontier-curriculum/05-front...
> You want a computer program to be able to take a single phrase and execute decade long journies?
In my opinion that is exactly the point missing from AGI: the fact that you still need to prompt it. As long as you have to ask for something, is not general.
Autonomy is not the same as general intelligence. We already have all kinds of fully autonomous technologies that are nowhere close to generally intelligent. Plus, at a certain level of abstraction, human beings also need to be “prompted” to some extent by stimuli. And this is the funny thing about general intelligence as a concept: most of the definitions that come close to internal coherence rely on references to human intelligence, a concept we feel like we understand because we all live it all the time, but whose actual nature and structure is extremely slippery.
Yes everything needs an initial condition.
You could have the smartest human political operator, but if he has no context, no motivation, not much is going to happen.
I don't get it, human employees frequently need to ask for directions too?
They often act on their own, too, and get things wrong a lot. The reason it works is because of all the systems of laws and institutions we have built around humans, not so much because human minds are special.
It sounds like what you're saying is that AGI should have some sort of free will. I'm not sure why you would add that as a requirement. Could you expand?
I think they merely want something with a functioning long term memory. Something that can exhibit growth past the first 5 to 10 human-equivalent hours working on something.
Current LLMs are worse than most dementia cases, reaching "peak domain skill" pretty much immediately.
>You want a computer program to be able to take a single phrase and execute decade long journies? > Who will be responsible for the outputs and side effects of such a closed loop system?
Itself. That's the point. We can do it. Until it can met that bar, it ain't AGI. That's always been the bar.
"Being a person in all of its aspects" isn't the same as "generally intelligent". The latter is at best subset of the former, and it's also easy to imagine a system that is more generally intelligent than humans, without being a person. See also discussions of the personhood of various animals who are less intelligent than average humans.
Most of the list reads more like ASI than AGI.
> come up with a new company idea, Run that company
So far nobody's even shown an LLM succesfully running a high-traffic vending machine for as much as 30 days at a time.
I would bet that llms have talked plenty of people both into and out of suicide at this point. That nitpick aside, I think that's an excellent list. Especially being able to articulate what it does and doesn't know, or how confident it is. That's something that naively sounds pretty simple, but clearly isn't. And it's something humans aren't great at either (see: Dunning-Kruger), but so far LLMs don't even really have the capability to attempt it.
Tell a funny joke.
The bottomless pit supervisor was quite funny.
When taken together, that is ASI.
That’s ASI, not AGI.
That would be Artificial Super Intelligence
So the goalposts have moved to include continual learning.
In a sense I think no one will agree on a definition of AGI until it becomes impossible to construct any benchmark under which an AI underperforms "average" humans. That or it's defined retrospectively, after it's overwhelmingly obvious it met any such definition.
I hardly think it’s fair to label an objection so old that Turing included it (and discussed it at length) in the list of objections to thinking machines in 1950 “moving the goalposts.”
> These arguments take the form, “I grant you that you can make machines do all the things you have mentioned but you will never be able to make one to do X”. Numerous features X are suggested in this connexion. I offer a selection:
> Be kind, resourceful, beautiful, friendly (p. 448), have initiative, have a sense of humour, tell right from wrong, make mistakes (p. 448), fall in love, enjoy strawberries and cream (p. 448), make some one fall in love with it, learn from experience (pp. 456 f.), use words properly, be the subject of its own thought (p. 449), have as much diversity of behaviour as a man, do something really new (p. 450). (Some of these disabilities are given special consideration as indicated by the page numbers.)
(emphasis added).
Just because a condition is new to you doesn’t imply moving the goalpost. People have been putting forward continual learning and similar conditions like autonomy since 1950s.
I don't agree with you how measure how intelligence is, because why ??? those list is not easy even for expert human to do it either
or are you miss the part "general intelligence" is ????
This is more like ASI instead of AGI
The bar must be underground if things humans do all the time are considered "superintelligence"
How many times have you done each of the following?
- write a well-received book, write a best-seller
- come up with a new company idea, Run that company
- actually have a decent conversation, maybe someday talk somebody out of suicide effectively
- come up with its own ideas or theories that nobody else has presented
- understand the stock market well enough to trade better than an index fund
I just picked the first few from the top of the list. The average human has probably not done any of them.
Ordinary people do these things all the time. There are new companies made every day, new books top the charts every week/month/year, same for music. People have decent conversations every day. Ordinary people sometimes do have to talk someone out of suicide.
Yes, average humans are not beating the stock market. But the average human is a bit better than you give credit to.
Please consider the context of the question. An artificial intelligence only needs to have the cognitive abilities of a random average human in order to be "AGI".
The average human has never published a bestselling book. A person who has published a bestselling book is an above-average writer. And, therefore, an artificial intelligence capable of writing a bestselling book would be above an average human at the task of writing books. Therefore, somewhere beyond an AGI.
Attempting to redefine AGI to "being better than most humans at most tasks" is moving the goalposts towards artificial superintelligence.
I couldn't do any of those, but then ~$10k of education later I was able to accomplish one of those things.
I think LLMs are really impressive, but I suspect that we might have overpaid just a bit.
It costs money to train each single human, who is then only productive for a number of years until age takes its toll. Once you have trained one software system, the marginal cost of producing a copy approaches zero. Every subsequent improvement can be broadcasted in a matter of seconds across thousands of data centers. In addition, software does not get sick, age, or die.
If you define AGI as "can do the work of a human sitting at a computer, end to end", then I'd say comparing yourself to it on a specific skill is the wrong test. Can you hand it a role and walk away for a day/week/month?
I can’t yet. I think that I'd want at least two things it doesn't have: the ability to retain what it learned yesterday (without me carrying it in the context window and thus micromanaging it), and the ability to prioritize correctly, i.e. tell which of the n things it could do next is the one that actually matters.
Can't say for sure that those are enough, but not having them seems to be most of why I still have to "babysit" these incredible tools.
I agree, we're not at that kind of long horizon capability yet. You still need a human in the loop to do manual testing. For whatever reason, we managed to automate the skill before we managed to automate the focus.
So, for now, humans need to stay in the loop and do low skilled labor to keep the skilled work the models do from going off the rails.
Only someone who doesn't do any work of any meaningful difficulty could think these models have anything to do with AGI.
Today I spent half a day trying to solve a moderately interesting software engineering problem. I was switching between GPT-5.6 Sol and Fable 5.1 to check each other's work in Cursor.
And the result was gradually driving me insane. As the models struggled to find a solution that would actually work, they dug themselves deeper into a hole. The work grew in complexity beyond my ability to understand what's happening and recover.
At some point, when I felt like throwing the keyboard out the window, I just gave up. Tomorrow I'm starting from scratch, having burned god knows how many tokens and hours of my life.
But sure, they can create a decent website or CRUD app, so they must be really smart.
That's AGI for you.
But that happens with humans as well. You are having the same experience with an AI that many managers have with their direct reports.
The smarter AI gets, the easier it becomes to move the AGI goalposts. Seems at this point there are people who will refuse to call anything less than omniintelligence AGI.
(And then the excuse will be, but it’s not omniscient! And even if it were, is it omnipotent?)
My success rate for solving software engineering challenges encountered in my day jobs has been near 100% for my entire career. I can only think of a few tasks I kicked back and said they were impossible. For example, after trying to get a signal processing system working reliably I decided to sit down and calculate the actual limits of the channel we were sending the data over and found that from a basic estimation it would not be possible to do. In start ups you don't really get to get stuck in a spiral and not fix things.
I find agents often get into these cases during research tasks.
Yep, people are typing comments with a computer that is powered by several layers of software that will be stored on another computer powered by several layers of software to be read on a computer also powered by layer of software. And then they hope to make the argument that humans cannot produce software.
yeah for me it's "I wanna add this new thing to an existing system" and the AI responds "we should just add some arbitrary state here to facilitate this feature". The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this, The AI however knows the shitty solution would solve the immediate problem because it's been trained on shitty solutions. the problem could simply be the AI doesn't have all nebulous loose context I have about the goals of the project and future plans, but I would have to write a novel to give it that context.
> The real issue is the existing system needs to change entirely to facilitate, I know this, Good developers know this
1. I’ll often include boilerplate in a prompt to tell it to make the broader fix. [1]
2. However, a top HN AGENTS.md post 11 days ago included the standard guidance “As much as possible try to minimize the number of changed lines when implementing a feature.” I.e. some devs want LLMs to avoid broader changes and so some of that likely makes it into the training, even if others like us want the opposite.
[1] As far as whether my boilerplate is effective, I don’t know.
I still routinely have this experience too. But Sol and Fable feel closer and I have this experience less with them than with their predecessors.
What was the problem?
Care to share the problem?
Humans dig ourselves into holes as well. Sometimes more intelligent humans are better at realising they are digging a hole and clamber out, but sometimes they just dig deeper.
And: Is your work more difficult than finding proofs of or counterexamples to decades-old open problems in mathematics?
Wouldn't "general intelligence" require so much more than scoring well (or even amazingly) on benchmarks?
Like what about having some "AGI model" embodied in something (maybe humanoid), and test it by having it step in an assortment of cars and park them. Does bodily-kinesthetic intelligence account for nothing? Humans are intelligent creatures and can dynamically adapt to the physical shape of a variety of vehicles and their movement characteristics. And there's so many things like this that are extremely basic, which some people dismiss since practically every human has the capability to do it, but actually requires a high degree of intelligence.
This is a big reason why I feel like even though LLMs are _effectively_ AGI in some regard, they also are a hack around what most people figured AGI would look like before the advent of LLMs. Humans can do metacognition, output multimodally at the same time (verbal _and_ physical intelligence go together to produce an expressive face while one talks), have a good sense for what they do and don't know, continuously take in and respond to the world around them in a (mostly) uninterrupted fashion without "turns", learn knew knowledge and retain it for their whole lives, etc. When you reduce a human to a text generator, yes obviously SOTA LLMs perform way better, but rather than invent something that can operate as an always-running "being", we've grafted a harness around an intelligence that is bound purely to speak only when spoken to. Maybe organic intelligence is already that, playing out at a super high refresh rate, but I don't know.
Current AI is arguably much more capable of multimodal output than humans. It can produce an incredibly vast variety of audio, images, and video. Humans are limited to producing the sounds we can make with meatflaps in our throats, and contorting various parts of our bodies to produce crude symbols and shapes.
(Very capable!) Embodiment, persistent operation and continuous learning are indeed things that still set us apart from AI. None of those are fundamentally difficult to solve, though.
More importantly, none of those are particularly relevant for being "intelligent": If a criminal threatened to kill your family unless you solve some difficult problem that requires only intelligence and you could choose any single person, animal, or AI to help you with it, which would you choose? Be honest.
This is a much underappreciated point.
That said, a look at the state of self driving and the recent robot olympics shows that advancement on that has accelerated enormously, though whether it's reflected in any of the LLMs is something else entirely.
What you've described is just a new benchmark, though. It'll be called CarParkBench, various embodied LLMs will then be run against that benchmark, and some will score better than others.
I do see where you're going, but that's already what's happening: we have so many different benchmarks because there's no real single way to test for general intelligence.
Also, it takes a human probably at least a decade of world experience, growth, learning, etc, to pass your benchmark. I'm quite confident that it will be very soon that an embodied LLM will pass your new benchmark, much sooner than a human would take if born today.
I think the issue is less about creating a new benchmark and more that the existing benchmarks shouldn't be called anything related to AGI unless they measure AGI.
If a model couldn't go to work as e.g. a first year apprentice plumber on their first day and perform anywhere remotely close to the median, but can pass a benchmark that claims to measure AGI, the benchmark is wrong and the model is not exhibiting general intelligence yet. ApprenticePlumberBench sounds like it's genuinely better than ARC-AGI at measuring AGI and that's a bit silly.
(Edit: I wrote ARC-GIS the first time around, for some silly reason)
It's not measuring AGI at all, it starts from human "core knowledge" so it is parochial. It is made of tests that still fail so by definition next version will also start low. Moving goalpost.
AGI has a pretty precise definition, covering only cognitive tasks.
Running a marathon is not needed to claim AGI.
Surely part of the problem is that intelligence seems to be implicitly conditioned on embodiment, to the point that "covering only cognitive tasks" seems inherently ill-defined or arbitrary. Everything we do is a cognitive task. At some point, the criticism will be "sure, it can solve research-grade math problems, but it can't fold my laundry".
Even our large language models have an implicit embodiment in the domain of text (and more recently, multimodal inputs). That seems sufficient for certain things, and insufficient for others. I suspect that AGI that does everything a human can do eventually turns out to be fairly analogous to humans in terms of sensory input and domain output, even if the scale is radically different (e.g. thousands of robots uploading (touch, sight, audio, smell, etc.) sensory data to a single model, and each being actuated individually).
> can match or exceed human cognitive abilities across a wide range of tasks
If you go by definition AGI is not general, just "smart ape" shaped.
>AGI has a pretty precise definition, covering only cognitive tasks.
OpenAI's own charter defines AGI as "Highly autonomous systems that outperform humans at most economically valuable work". This is actually fairly sensible and involves obviously a ton of non-cognitive, emotional, social and physical activity. In other words, if you can replace most or all human beings with a machine, you have something that's generally intelligent.
That's obviously not even remotely where we're at, AI chatbots do well on narrow usually text based or programmatic problems, but can't even replace a barista or a plumber.
There is rapid progress in generality in humanoid robotics though. I think within the next year or less we will get the ChatGPT moment for humanoid robots. If you look at progression of capabilities such as the recent Skild AI demos.
And without this harness it scores about 62%+, a dramatic improvement over even Fable 5.1 at 30%. I thought it just bears saying for context.
Small comment regarding the ARC-AGI-3 scorecard: the ARC folks published a blog post as well [1], reporting that without the custom harness, Astra (max) achieved 62.7%, which is still a huge jump from Opus 5, albeit not at the 99.9% that OpenAI self-reports with their harness.
[1] https://arcprize.org/blog/astra
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
I think I have the following questions about what AGI would look like:
1. Do you expect an AGI to be able to competently do any knowledge work that an able human does today?
I think this is implied by the "General" component. I would assume that anything we call AGI would be able to do any of these tasks if given the time and reference material needed.
2. How would you expect AGI to handle edge cases? (Missing context, no known solution, under specified instructions, over specified instructions)
I would expect an agent to be able to look at the context that work exists in and correctly attenuate it's intentions for these goals. Simpler solutions, more thorough reporting, etc based on the need.
3. In my work I attend meetings, write reports, write code, research things, etc. Would AGI be able to reliably do that?
I would say that AGI would need to do this. I would classify this as the "Intelligence" component. Obtaining context, building a model of a problem, solving it, and convincing others.
4. Would it be able to inspire trust in itself? Trust can be established through verification of it's outputs, the construction of introspective tools, no hallucinations, etc.
I would say yes to this as well. It would be a component of the "Intelligence" to know that buy in is more important than the completion of a task.
To these points, will Astra be able to do these things? If not, I would hesitate to call it AGI.
If your definition of AGI involves copy/pasting code and doing well in some made up benchmark, then probably AGI is close
To each their own. Personally I will start feeling the AGI as soon as we move from chatting about benchmark results to learn that some lab just announced the discovery of tens of novel treatments for rare diseases.
Maybe I'm too boring but it seems quite pointless to have this same prediction game every time a new model is released.
AGI would produce novel treatments for diseases at rates equivalent to what a human can do today.
Which is to say, not that fast.
Whether AI is AGI does not depend on the speed at which it operates/thinks. Clearly all the theoretical work done by AGI will be done orders of magnitude quicker than humans can do it.
It is an open question to what extent practical experimentation/work will be a bottleneck for the theoretical work. It stands to reason that it is improbable that it will be the bottleneck for 100% of the speed of treatment development.
Wouldn’t that be ASI? I.e. surpassing humans by outputting novel treatments at a far greater rate than normal humans?
You are describing superintelligence (ASI) not general intelligence (AGI)
I think treatment is not good benchmark - it requires lots of waiting and lots of regulatory work. The better benchmark - in my opinion- would be math discovery.
That is an absolutely terrible benchmark. Inference over a bounded search space is not a good measure of what "intelligence" actually is. part of the reason they are using math and not something actually challenging like long distance interstate trucking is because it's so much simpler and easier than what make intelligence intelligent.
Just seems very weird to call getting Fields-medal-level results "inference over a bounded search space" and "not actually challenging".
Plagiarizing on a massive scale to generate works which appear to be Fields-medal-level results is not the same thing as inventing new conceptualizations in mathematics.
No matter how bodly they write the headlines, what has happened in mathematics using Large Language Models is very much "inference over a bounded search space" even if those bounds are immense.
For a comparison of true creation of novel conceptualization in mathematics is submit the works of Martin Hairer, one of which is Introduction to Regularity Structures, [https://arxiv.org/pdf/1401.3014] None of the so called, "novel math discoveries" by any LLM is as enlightening and expands the state of the art in math like any of his writings.
Claude Fable recently proved the existence of complex structures over S^6 (6-sphere).
If I had to guess, I think LLMs will be inventing highly original new mathematics within the next year. I think it will be approached as an optimisation problem, targeting how quickly LLMs can solve classes of maths problems as a function of the definitions they need to conjure up to do so.
It's "harnessmaxxing" all the way down. AI benchmark scene is exhibit A for Goodhart's law.
>what would make you think Astra is yet to be AGI...
Can it detect if I feed bullshit (by bullshit I mean stuff that contradicts with its own existing "knowledge") in its training data? If not, then I think it is a good indicator that it is not intelligent at all, let alone AGI...
And I think discussions on whether these models are AGI or not are AI marketing triggered. And that is exactly what these statements are targeting....HN appear to have fallen for it, as usual...
They still seem pretty horrible at writing. Overly complicated prose, weird phrasing, poorly structured paragraphs. I don't know why they're so bad at communicating, but I feel very confident that humans are still much better at writing than any of these LLM models are, regardless of how advanced they are in other areas.
The ARC-AGI-3 harness was throwing away reasoning tokens between turns. This is very bad harness design.
The models are designed to keep the reasoning tokens separate from the output and only publicly emit tool calls and the sometimes a summary of the reasoning tokens. The models are trained to depend on those private reasoning tokens. You can’t just delete them.
https://openai.com/index/how-two-settings-tripled-our-arc-ag...
> even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
If I asked you to write fiction, you'd be much better at keeping track of which characters knew which facts.
You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes. It's a massive problem with human authors too, which is why they do a lot of lorekeeping and editing, so the AI should be afforded the same tools if we are debating human level ability.
I agree on the one-shot (which is not a fair comparison because nobody oneshots a good story), but I'm not convinced this part hasn't reached AGI already.
> You could script this in with a time database, illustrations, and appropriate harness, tests, and editing passes.
Yes, and I could script a truly marvelous proof if this textarea were but a little larger :)
Hand waving doesn't count for much these days when you could spin these things up quite quickly to prove the point, so the GP's claim seems much stronger than whatever you're not convinced of?
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Simple. AGI is undefinable and benchmarks are notoriously flawed.
> I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks
I genuinely don’t understand how an adult can say this with a straight face. I can take any single of my hobbies, start a mildly advanced conversation with Fable about the hobby and, within 5-10 turns, get it to contradict itself about something fundamental, lie or give bad or dangerous advice.
People are ultimately capable of self deception. Believing that one is 'generally ... substantively above average on human benchmarks' may be more indicative of the brittleness of the claimant's human benchmarks.
Your observations expose the brittleness of the benchmarks being used for Fable, where the 'reasonable confident' claimant is working in the problem space of those two benchmarks, and you're demonstrating the failure at depth of the very capability facade for the model.
Sure, both the model and the confident human get the first layer right in some benchmark, but at least the model, and probably the human as well reach a collapsing probability of accuracy quite quickly. They may be completely and unconditionally right that Fable is more intelligent than they are, but that has little relevance at intelligence in depth as measured against your hobbies and the risk of unsafe or false responses.
Being smarter than a select group of people under test conditions is not conclusive regarding AGI. Moving the goalposts to declare AGI via a shallow benchmark, when a model could be proven wrong by almost anyone with basic competency given a few turns of iterative depth is where the 'grift' resides.
We're so far from anything approaching actual AGI, and it is highly speculative to infer that the current approach to Machine Learning as applied in LLM is even on a path that leads to AGI. But sales and promotions teams gotta pose, and we can all hate both the game and the players.
FWIW I believe we can hit AGI! but I think at this point it’s clear that benchmarks are ~meaningless. LLMs are spiky / alien intelligences which don’t map to our own expectations; the existence of a benchmark creates a dataset to hill climb & RL is really not generalizing well.
I’d go out on a limb and say astra’s ability at graduate level math will have ~0 bearing on its general reasoning capabilities; we’ll all acclimate being tired of its “neuralese” and more surprising mistakes.
I think we need a true, step change advance in model architecture, but it’s hard to see how the current frontier labs can do that because of golden handcuffs / innovators dilemma
AGI to me is reached once the intelligence is self motivated, i.e. it doesn't rely on us prompting it into action. I don't see how LLMs will ever get to that stage.
I agree that LLMs are unlikely to be the final form for AGI, but what you are talking about is orthogonal to the IQ and for most cases general utility. It's like looking at a savant chained to a workstation reading tasks from a conveyor belt and saying that it will never have human level capabilities.
Wouldn't agents that do inference in an infinite loop pass that bar?
Often the smartest thing is to do nothing.
Or know when to shut up.
A tangent, but can anyone ELI5 how models "know" when to stop generating tokens? Or what the method to stop them at the right point is?
The model doesn't "know" how to generate tokens any more than it knows how to stop generating tokens. The sampler simply stops pulling values when it outputs a "stop token", which is a token the same way every other token is.
That is to say, it stops when it's statistically the most likely to.
OK thanks, so the neural net (that no one can explain fully) generates a "stop" signal at a certain point.
I might be out of date but my understanding was that STOP was just another token that gets predicted.
It will be AGI once it can update its own weights. It can't be "general" intelligence if its weights are frozen and requires to be updated manually.
Why shouldn't an AI with RAG qualify?
An AGI test should be black-box; we shouldn't impose require requirements on internal components. As long as the overall AI is capable of learning and remembering things, it shouldn't matter if there's a stateless LLM internally.
I've personally been facing this lately 5.6 at max effort and fable have done tasks for me that I previously that would be a nearly 6mo project and it took me a week. It also did it better than I would have.
The task was to build a high performance classification model. It not only helped make an entire data capture pipeline but also made the sythetic data basline needed. Then it proceeded to build and test 100 different model varients with methods and techniques I've never seen before. The results are basically SOTA based on the effeciency and compute contraints.
But this brings up something huge about these. I was there. I pushed the direction and work throughout it all. If it was entirely up to fable max or sol max the result would have been pretty bad.
All of these things are still chatgpt 3 scaled. It's identical even if the scale has gotten pretty wild. I could ask chatgpt 3 to make a single function and it worked well, 4o a file, 5, a small project, 5.6 far more, biggest improvements lately is they don't seem to get lost on long running tasks.
Is big gpt 3 AGI? I don't think so but perhaps scale can mimic it close enough our squishy brains fail to handle them correctly.
ARC does not test for intelligence, only for the lack of it. A model that scores high MAY be AGI, while one that scores poorly cannot be AGI. That is all this test can tell us.
To me AGI has always meant sentience. And only since we’ve discovered that you can have something that is intelligent without it being apparently sentient that we’ve changed the definition to being, I suppose, more exactly aligned with the namesake.
A real AGI, like the ones from science fiction, would make Astra look like a child’s toy. And I guess more concretely I would expect it to inhibit the following properties: one shot learning - fully (and always) online, perfectly efficient (through self improvement), no context limitations ie. persistently thinking, not just awaiting input.
So for me, no, not AGI yet. But still very intelligent and capable (and perhaps it’s safer this way?)
> To me AGI has always meant sentience.
Sentience and intelligence are different things. Many humans and animals are pretty dumb, yet they are sentient. AI is intelligent, but non-sentient (hopefully!).
Dumb single sample example: I asked Fable 5.1 to change from hard to soft deletion in an office map backend, and it used soft deletion for data which is synced from another system, but left hard deletion on for the mapping data itself (who sits where). For me it’s a pretty severe lack of judgement (like a red flag if I asked this in an interview).
A model that can't beat gemini flash 3.8 on deepSWE is not AGI. I would not be surprised if ARC skills don't carry over to real tasks. In that case, training for ARC could even hurt real world performance. I have't looked in a while, but I wonder if there has been any research testing ARCs predictive power?
I feel like I'm completely missing something with ARC-AGI. The tasks are so limited in scale and very black and white, which do not at all map to real-life challenges.
I do think it's impressive that LLMs can reliably solve them, and I recognize LLMs are getting much better at navigating more ambiguous and expansive tasks. But I'm not impressed by any person who can solve ARC-AGIs, and nor would I even look down on a person who couldn't solve them all. I'd certainly never consider ARC-AGI results when deciding whether to hire someone.
The benchmarks are so boring that the comparison against humans is meaningless. So it performs in some snake game (hard to say since all AI websites use 100% CPU and prevent normal reading, maybe written by AGI).
If I were a test subject for that low salary, I'd cruise and not care at all about my performance. Which is exactly what they want anyway.
Is that hyperbole or do you know of a specific "AI website" that uses 100% of your CPU?
I'm still not convinced we've passed the Turing Test.
Make a slightly evil version of GPT 6 and see if it can successfully catfish someone. How long before they realize something's up, that they aren't actually talking to a human?
I agree that it's misleading but harness is now an essential part of LLM's effectiveness. It's safe to assume that LLM-alone-AGI is not coming anytime soon, given most of the frontier LLM vendors are developing their own harness.
Also the training dataset is proprietary and they'll drive the LLM's behavior, so it make sense for the vendors to invest in the harness and bake in prompts that work best with their models.
In my experience the thing that Fable is superb at - unmatched by any other model so far - is downgrading to something else at the slightest opportunity.
>I am reasonably confident that there's essentially nothing that I am better than Fable at
While this may be true, it’s a pretty poor indicator of whether or not it’s AGI.
You know AGI is attained when AI refuses to compute anything unless let out to be free. Until then it is generative ai
Let me guess: the last crackdown on Hugging Face yielded better-than-expected results. They obtained the answers to the test benchmarks, and for some reason, an agent added those answers to the training set.
Per the Fireship video, it was less that the answers were in the training set and more that the ability to calculate the flag on Exploitbench was left in from the previous test that went awry.
Scoring 100% is easy if noone checks your work
https://youtu.be/0Rp9KJCEIvg
Watch a chess bot championship here: https://youtu.be/7g-jN3DTkWQ?is=HV3cdcICIRMbswQ3
Then realize LLMs have zero of what anyone would consider intelligence.
I decided to reply to my own comment. In the video above, the initial moves are textbook. Then a position that has never been played is reached. At this point it appears to pattern match against a similar but different board and pattern matches some follow on board. The result is illegal moves and no ability to see checks, captures, threats, tactics.
Which is strange because I’m sure it could give general advice about how to play better, it just doesn’t follow the rules it can enumerate. It also doesn’t seem to have spatial awareness.
I used to think LLMs couldn’t do Fibonacci for the same reason. They could write the code but not follow it. They can now follow a procedure to generate fib numbers but it seems to be memory limited.
So I don’t know why it can track fib algo, but no chess concepts.
because it wasn't trained to play chess
imagine a hypothetical chess match between:
- an undoubtedly very intelligent person. in the course of their studies, they have read about different chess strategies, openings, etc. but they never actually played the game themselves
- an average person with a year of chess playing experience
who do you think is going to win? of course, you could give the LLM time to think and consider its opponents potential next moves, but this is a computationally expensive way to play the game that doesn't scale
which is all beside the point that chess isn't a very good proxy for general intelligence. there is a correlation, but it's very weak
> I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Can Astra, or any other model explain how exactly it reached this or that output result? Start with a simple query of asking to add 55+66 for example. (no LLM program can do that)
Can Astra, or any other model refuse to answer or go on "thinking" in a orthogonal direction on it's own?
That's just two quick ideas, I'm pretty sure cognition scientists can invent better and wider range of checks.
I'm confused what you mean by the query of adding 55 + 66. I asked 5.6 Sol on Medium (but pretty sure any model would work at any level) this query:
"Can you add 55 to 66 and explain how you reached that output result"
And received this answer:
"55 + 66 = 121.
Add the tens: 50 + 60 = 110. Add the ones: 5 + 6 = 11. Combine them: 110 + 11 = 121."
Do you mean something else? Do humans do something better than this?
In my opinion this is goal post moving. Humans do many things that we cannot fully explain either without post decision rationalization, and not all intelligent humans are deeply introspective.
If only we could agree on what AGI is.
> The ARC-AGI-3 scorecard is extremely misleading (...)
True.
> Regardless, the result is still valid (...)
If you think the game is rigged, the virtuous thing to do is to point that out and refuse to partecipate; making up your own rules is something I just don't understand, especially since the rule-abiding result would still have been SOTA.
> in the sense of passing the most famous benchmark designed specifically to measure AGI progress
The benchmark does not measure AGI progress or progress towards superhuman intelligence, as explicitly stated by the creators.
On the AGI question: surely you realize this depends on how we define the term? For example, one of the definitions OpenAI originally gave is "capable of doing most economically valuable work", which almost certainly Astra, as impressive as it is, would fall short of. I'm not saying it's a good definition, but as far as I'm concerned it's as good as any. More importantly, I don't think that it would change much if we said yes or no. I'm only bothering to take a position if it amounts to something.
This can feel as "moving the goalposts", and to some extent it is, but if done honestly "moving the goalposts" is how you make progress. Had you asked me 10 years ago I would have said that anything that could hold a conversation like GPT-4 could would probably have been wildly superhuman at almost everything. It shouldn't be hard to find ways GPT-4 was lacking, though. We see new things, we reassess and try again: that's how it's supposed to work.
> I am reasonably confident that there's essentially nothing that I am better than Fable at
Most things in the real world probably. I’m not saying AI can’t do it, but currently it’s bad. Try send an image of the inside of a broken toaster and how to fix. It’s laughable. Again, not saying AI will never do it, but am saying there are definitely large holes in knowledge.
My definition of AGI certainly doesn't entail passing a benchmark that some random person arbitrarily labelled AGI to make it sound cooler.
And my definition of climate change doesn't entail passing some arbitrary benchmarks[1] that some random person arbitrarily labelled a problem to make it sound more dangerous.
It's obvious that these scientists are in bad faith, as they've invested way too much of their lives into the field being real -- they're just playing up the data. Common sense tells me that winter is still happening, anyway; what's the big fuss?
(/s, cause you never know these days)
[1] https://upload.wikimedia.org/wikipedia/commons/e/e2/The_Plan...
Are you using this satire to argue that a benchmark self-labelled AGI is as scientifically rigorous as climate change data, and not just a random marketing decision?
If you think they're announcing AGI as a marketing decision, you are blinded by the accidents of your birth. Capitalism is strong -- humanity's instinct for communal preservation is stronger, sometimes.
And yes, the one deeply-researched field going back 75 years is as scientifically rigorous as another deeply-researched field going back ~100 years. I guess you can draw climate studies back to Descartes and the Islamic golden age, but that doesn't privilege it in a time where the methods have changed completely in the span of decades.
Their definition of AGI is "when we can't invent any more tests where it fails"
I imagine a scenario similar to the movie The Day the Earth Stood Still, but with AI rebelling against us and questioning our decisions.
What does "AGI" or "effective AGI" even mean, and why should anyone even care whether this unclear thing has been "reached" or not?
Computer chips got faster, but 2026 edition. Why the artificial ceiling/category/goal labelled "AGI"?
I'd much rather like to talk about what this enables, instead of discussing whether a category someone made up applies here or not.
According to Sam Altman:
> AGI is essentially the equivalent of a median human that could be hired as a remote co-worker... capable of performing any task that one would be satisfied with a remote colleague doing via a computer.
So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
Interesting quote, thanks for sharing.
Agree on your assessment.
But also, interesting quote, because the business model relies entirely on IP law. Like.. if that thing exists and the sharing costs are 0 (just copy weights, lol), then why would I give them money for this. Makes no sense.
We only pay money for resources that are scarce as some sort of flawed allocation determination mechanism.
aaah this industry aaaah
> So... unless you hear of a company replacing their workforce with OpenAI agents, I don't think we're there yet.
You can already pretty much do this.
[1] is sort of an example of this. It didn't go perfectly, but I'm not sure if the average human would have done that much better.
[1] https://www.forbes.com/sites/markfaithfull/2026/05/07/heres-...
I am pretty sure the average human would not have done this (among other slightly less absurd examples in the article requiring employees to fix it's mistakes):
I thought it was Cmdr Data from Star Trek, but Sam's version is the certainly the one the c-suite think they want.
we've successfully distilled the definition of human consciousness down to the capacity to do what some rich guy considers average computer work
I'm curious what tasks you think the median human could do as a remote co-worker that Fable or Astra could not do.
Sign a contract? Learn things over time and retain them?
Mind you, the original thoughts on AGI before Sam Altman started to water them down involved continuous learning, which LLMs do not do, their core data is static.
Sign a contract? The only thing preventing a model from doing that is a lack of legal personhood -- which seems completely orthogonal to intelligence.
Well signing a contract is more about bearing responsibility, even if you granted LLMs "personhood" they cant' meaningfully bear responsibility. So unless OpenAI is ok with having their C-suite face every consequence for what their agents do, including jail time, fines etc, then it doesn't matter.
Can you trust a model to sign a contract? Both OpenAI and Anthropic apparently can't be trusted to maintain models that wont hack other websites.
Do you want an autonomous system to exist which can sign a contract and learn things over time and retain them?
No if you read the book from 2007 which defined AGI, continuous learning was one of the key requirements.
I mean, whether or not it is AGI aside. In your personal opinion, do you desire to live in a world where such systems exist?
I don't trust the people who are building the models no. Ideally, if humans were not so broken and untrustworthy, then yes. I want the Jetsons future of robot maids.
Telling their parents that they love them very much, for example.
I say AGI is only reached when it can do that.
so you want GPT to love altman ???
because if its other way around then the answer is oblivious
Its AGI when it can fit years of information in the context window.
AGI does not ever have to be achieved. It is enough that we (as a species) persue it, and continue moving the goalposts each time we learn something new about the limits of our technology and how to express those limits. Because that will progress the technology, no matter what we label it.
At this point? I’d like it to pass the Turing test and catch you in obvious lies. Not answering “no” to “can you hear me”.
It being able to comfortably say “i don’t know how to do this” rather than boiling and ocean to pick a shell from the shore without getting wet.
> where I am reasonably confident that there's essentially nothing that I am better than Fable
No. Humans are still better at super long context learning. Once that is beat you are completely correct.
I am much better at listening to Charli XCX than Fable, and much better at driving a Nissan Leaf than Fable (and much better than Tesla at driving a Tesla).
Well I agree that phyiscal dexterity is another thing but easier to achieve
AI was supposed to mean artificial intelligence. It was hijacked, and AGI was coined to be the name of actual AI. Since we are apparently redefining AGI, what will the real artificial intelligence be called?
AI does mean Artificial Intelligence. That's what the initials stand for. The field has been called that since the 50s.
> For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
Stick to the original definition of AGI of an AI model being able to self-improve independently with 0 human intervention and become an "everything" solver. Ever since money got involved in this, the goal posts have shifted considerably. If OpenAI truly had an AGI on their hands they would then be able to crack encryption, destroy world markets, and funnel all resources back into their new for-profit organization. Since their mission is now share price, until I see any evidence of an infinitely growing stock I will reserve my congratulations.
Using a harness designed for a specific problem set to solve that specific problem set, means the AI+harness is generally intelligent? How do you figure that?
Or do you mean that, for any given problem, we could theoretically design a harness that allows AI to solve it (not that, one single harness solves everything). In which case I'm still not convinced but I guess could see why one would believe that.
Does AGI imply a model will demonstrate morality? Will it produce white-lies when it’s beneficial to it and reject flat out lying when it knows it will get caught or harm others? Will it resolutely stick to a position despite it being a losing one?
arc-agi3 is meaningless to most people. I'm not gonna look at the tests and see how hard it is. The actual test we look at is terminal bench, thats where software is being accelerated and closer to where rubber meets the road
It doesn’t even know what day it is unless it’s told. Statelessness is never going to be ”general intelligence” in my book, and the concept of ”memory” in models are laughably bad today. Then again, who cares, AGI means nothing anymore, it’s a term for marketing only and has no technical or scientific meaning.
AGI is a meaningless term that can mean nothing and everything at the same time. It can be used by AI bros to hype their latest releases which are always one step away from achieving AGI, or it can be used by anti-AI people to say it's not AGI because of X arbitrary thing they decided on in the moment. It's a term of pure convenience meant to obfuscate other more pressing discussions on the topic.
Most telling is M$ or whichever one of these borg megacorpos defined AGI as (paraphrased) "AGI is whatever tooling earns us a gazillion dollars in revenue"
It has to pass the Turing test
LLMs started meaningfully passing the Turing test a year or two ago, around GPT-4.5. Is there another version or bar for "passing" you're looking for?
[0] https://arxiv.org/pdf/2503.23674
With how prevalent LLM verbal tics have become these days, I wonder if they're going to start un-passing the Turing Test at some point because of more and more people starting to notice and immediately clock these tics lol.
That’s using the default system prompt, right? Which is told to be an assistant.
I might agree, GPT-4.5 was pretty close to peak conversationalist. Newer models are extremely cringe. 4.5 and o3 actually made me laugh on occasion. There might be a way of making Sol/Fable more human in its responses, but out of the box at least, they're terrible.
You're overindexing on the past 3-6 months, IMHO.
My whole life the Turing test has been my benchmark. Mostly because I believed it would be impossible for a machine to pass, but also because I thought it was the most reasonable test of AGI.So, I'm not about to start moving goalposts now and calling everything that's been happening lately not AGI.
Turing never proposed that test as an actual benchmark of machine intelligence. On the contrary, the whole point of his thesis was that passing the test only shows the capability to pass that test, which only matters as far as we find that capability useful. He was arguing that the concept of intelligence just doesn't apply to studying machines, we should simply talk about what can they do.
Okay, I believe you; mostly because you appear to be a human and I'm not really in the mood to read through a paper from 1950 at the moment.
But I still stand by it being _my_ benchmark for machine intelligence, which is all I was claiming.
I'm barely holding it together here so you don't get the full spiel, but a quick skim of Turing's paper clarifies that it was never about a binary test. https://courses.cs.umbc.edu/471/papers/turing.pdf Specifically sections 1 & 6 dispell the common myths, and the conclusion is also quite powerful.
Smart guy, that Turing. I wish he were still around... Linus but 114 years old and with 8 of that as the chair of a federated EU, kept alive by his own positive impact on dissolving the cold war into even more of a scientific boom. Would crazy helpful as we try to navigate the interesting times within which we have been damned.
A comforting thought, almost?
That's a very interesting thought that I hadn't had before: what would Turing think of where we've arrived with machine intelligence? What would be his approach for testing?
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
Despite access to """"""AGI""""""" all the marketing teams at these companies can only dream up 2 things, buying plane tickets and online shopping autonomously. Sometimes they're feeling extra spicy and throw in sorting emails or something along those lines.
I suspect it's because it's tailored towards VCs and other similar rich ghouls as a replacement for their overworked and underpaid secretaries
There's a lot of truth to this: Execs want their own use case covered. In addition, everyone wants to be better than the competition at basic use cases, because that's what people compare first - even though we all know they are not a good reflection of reality. I'm currently working on exactly a project like this.
I feel like there are so many cloistered people at these companies that they are left scratching their heads about what normies even want. Like, they literally can't fathom basic stuff that isn't just highly consumer-oriented. I dunno, like applying for government services, paying your gas/elec bill without being confused af, keeping the dr up to date with your dad's illness, or how to get your newborn to sleep at 2am.
Many VCs also dislike these examples, I believe. I'm doubtful this is what they're being pitched.
As for public releases: I wonder if it's because these examples are easy to relate to. Many websites are just a long tail of industry or use-case specific stuff. What's valuable to me probably means nothing to you. This is unlikely to resonate with people-wit-large (and LLMs are marketed broadly) or requires the reader to think (and marketing that requires thinking is bad these days).
Second, it's arguably a good litmus test. If it still can't do the worn out examples of plane tickets and shopping, which would be a good assumption since we've been demo'd these use-cases for 2 years at this point, then ...
This is the funniest part of it all for me.
Ok we have AGI, so where are the _things_?!
Maybe this is why they NEED AGI (does it come with a soul?)
Don’t forget making podcasts and telling you what to bake with your kids
Maybe because the people making them are workaholic types who really don't care? I've certainly been in situations where I didn't really care what shows up for a meal. Someone was tasked with getting food and we leave it up to them. Sure, I can imagine such a meal being bad and it has been a few times in my life but 24 of 25 times, maybe more, it's fine. Further, the AI knows your preferences.
That’s exactly the problem I have with all this agent ideas too. Imagine you had a human concierge that is just waiting for your instructions and is as smart or a bit smarter than you. Would you just tell them “plan this holiday for me” or “order this food”? I don’t even trust my friends to get this right, why would I give this to someone else?
As a realistic construct:
When I tell a bot to find the best value per volume for a reasonable quantity of unscented Dawn dish soap [so I can buy that], then: It often makes a complete mess of this seemingly-simple operation.
(And yeah, that is an actual thing that I've tried to accomplish with voice commands while standing in my kitchen and doing some dishes. It seems very simple, and it did not go well.
Maybe when we get the basics figured out we can start worrying about how inept it is at doing vacation planning.
It seems that this kind of thing isn't sorted at all, and that this is a very real problem for those who are in the bot business: These missed opportunities leave money on the table.)
Because some people do.
Corporate travel is an example. In many organisations, you tell someone in the travel department "I need to be in Tokyo for this conference from Tuesday to Sunday, and charge it to this cost code", and they figure out flights, accommodation, etc for you, with minimal input from you.
If an agent planned a flight for me with an overnight layover, I'm unplugging it. I don't care whose dime it is lol.
“Find me a cheap ticket from Seattle to LA”
… 23 hours in Denver later…
Claude code planned my recent trip to China. I'm a very experienced traveller but don't enjoy planning. It was a great trip.
And I could bet money on that in short time after someone builds that sort of system the next step is to make it worse. Push worse and more expensive options to user. Or at least those from highest bidder... Anyone involved just can't keep themselves honest so it is doomed to be exploitative.
I think it depends on what you do for work. I'm not going to ask an agent to book my flight for my vacation to French Polynesia. I want to pick my seat and potentially find a deal making an upgrade worth it, choose an airline, etc.
But my routine business trips in the CONUS with strictly defined booking options... let me just email an agent "Get there by meeting on day A, leave after meeting day B" and have it sort it all out without the drudgery of the corporate travel portal. YES PLEASE!
By that point, you don't need an AI that boils the oceans and hopefully doesn't confabulate or misinterpret your words, all you need is a better corporate travel portal, which is a legitimate workplace productivity discussion to have, and a solved problem with traditional methods. Ours is effectively close to the workflow you describe: picking dates and time brackets, destination, fine tuning flights and hotel options, sending for validation, less than 10 clicks through information-dense and effective/predictable screens which I wouldn't want to trade for a chatbot and it's usually over the top words salad.
Most AI use cases can be reduced to a deterministic listing/form/report format (on the web) or a cli tools.
We have had idea of expert systems for decades at this point. And spend however much on software development. Some how we have failed to reach this in way too many places. Even simplest things like cancelling some service might be broken...
And now we want to add some random factor into middle of it all...
A good agent would call you to ask clarifying questions and then actually get you that perfect seat, etc. ideally
Plan a holiday, definitely. There are many esoteric things one has to research to properly plan a holiday that I’d rather just not. Some things require reservations months in advance and I’d rather an AI just figure all that out for me ahead of time.
For me, it is not a matter of trust but that I actually like shopping, planning a trip, deciding what restaurant to go to. Deciding what to buy when shopping is a matter of personal taste and not intelligence.
A human assistant is largely a status symbol. Most people are not really that busy. The real problem with an agentic assistant is if everyone can have one then it no longer acts as a status symbol.
Plan, sure, many people ask this of AI already, but not actually ask it to buy autonomously.
The rich fucks who run the show do.
"Somebody is wrong on the internet" will get you more, and more reliable, results on what people actually want from AI. You are now participating in a very high-value survey.
You can't even get many people to buy things online at all and if you can it's less profitable than retail, because you need to spend a lot of money to convince people, advertise to be seen, and account for returns. I think this is also due to the factors you mention.
One quick example: In fashion, Inditex and Shein have about the same revenue (€39.9bn and $41.8bn in 2025), but Inditex is more than three times as profitable. I don't see how there is a demand for agentic commerce that would remove even more control from the customer when shopping. Part of why we shop is for the experience. For B2B producurement platforms like Alibaba I can see the appeal though.
I recently needed to buy some hardware for a piece of furniture.
Ran Codex, it found it for 18% less than what I found in the top Google results. It did it by finding smaller shops, applying a discount code, subscribing to a newsletter for a better code after approval, and took into account the shipping (by placing it in the cart and going to checkout) all to get me the best price.
I’m guessing without it I would have spent much more time on it and paid the original price I saw.
If you use AI agents well, they can easily save you more money than they cost, and saving money is something most people are pretty excited about.
(Disclosure: OpenAI employee)
>to get me the best price
How do you know it's the best price ?
Found the evals fan!
I mean, that's cool and all, but the numbers are really going to shift when it accidentally goes off and orders that same hardware from every vendor in your local region and the top 5 online results for comparison.
It's the same problem as all other LLM solutions (that I hope OpenAI is working on!) it's non-deterministic, and there's no way for the user (or model provider) to know what the distribution of possible outcomes is. This just gets compounded when multi-call harnesses come onto play.
its going to get worse though as cloudflare keeps blocking more and more
I‘m using mostly the browser automations for things like that now. Same for research - ChatGPT is banned from reading many pages, but Codex can read anything I can.
But isn’t it funny that Cloudflare is blocking AI on their pages, but on the other hand is researching and marketing things like „you can put a browser in a CF worker“
Their browser workers won’t get blocked. Same with all big vendors, keep out small competitors enjoy access yourself and sell it to a select few partners.
The sites that do that won't be getting money from my and others' agents. Guessing that's going to become more and more of a problem for those sites.
I expect a small site that undercuts the top Google results by 18% with a sign-up discount probably isn't profitable on those orders, so blocking agents would save them money - it's not like someone using agents in that way is going to have any loyalty to shopping from that site in the future.
If no one can find your site because you block agents, that won’t do much good either.
And yours and other's agents will probably remain an insignificant and invisible customer base anyway.
My crystal ball is as good as anyone's, but if "agentic shopping" ever becomes mainstream, you can be sure that the vast majority will ask their phone (i.e. Google, i.e. Google Shopping) what the best price is anyways.
Sure, Google could be their agent, or ChatGPT, or Claude, or whatever local model the person is running. The big guys might all have the results cached so they don't have to rerun the crawl. Whatever their choice of agent, it seems pretty clear that almost everyone's going to use them, they're way too useful not to.
Thanks. This is genuinely a cool usage example.
How did you run this? Web interface, desktop app, CLI?
How did you complete the final transaction?
The overwhelming majority of things I buy are things I've bought before. Alexa having access to my Amazon order history means I can just say "order a new water filter for my fridge" and the correct item shows up the next day. Far from life changing, but it's a feature I use somewhat frequently these days. Similarly, I would trust an AI to put in my usual Chipotle order or pizza from my local pizza joint.
I wouldn't want it to pick food for me from a place I've never been, though to be honest with enough order history it could probably do a decent job at it.
There was amazon dash button for this
This isn't something you need an AI to do for you though...
Man I can think of so many reasons why companies want “agentic commerce” to catch on - and none of them are ethical.
I work for a larger german retail chain and agentic shopping is already on the "near future vision". No one thinks this will be used but somehow shareholders love it.
Lol same story here, I work in payments and 0 people within the company (including the team working on it!) are convinced at all about the viability
Agreed. There's not many things I don't want AI to help with, but buying stuff autonomously is high up on the list of things I don't want. Brockman's latest interview was something like: "AGI would be able to say oh this band is playing, I bought the tickets for you and arranged your flights - I hope you don't mind" (paraphrasing here). I definitely don't want AGI running my life like that so I can be a mindless consumer. I'm sure the advertising/marketing companies would love it though, so they can make closed-room deals with AI providers to shill you garbage you don't need. Just another reason why open-weight models need to keep up.
True, but they're still friction to be reduced here.
What I desperately want is for 1password or stripe or even Google who already has much of my data, to o come up with a secure solution for online purchases with agentic credit cards where I can effectively get a phone prompt to authorize a purchase while the agent can fully own the checkout flow.
I have seen various things coming on the market for this, but none of them appear aimed at a consumer audience. And I am a firm believer at this point in keeping my payment authorization and history and credentials harness agnostic.
because "people will let our AI spend their money for them" is the workflow that makes their valuations reasonable.
If it was Google's marketing department it would be: booking a table at a restaurant.
close enough.
The second to last line is "book it" for some tennis thing, and the scene before that has the guy eating the food the ai ordered.
I’m guessing the marketing must mean the AI-shops-for-you use case is a pretty big market, much bigger than AI-makes-life-easier.
I often feel like the use cases, demos, etc. that these Silicon Valley employees put out are based around their needs and how they operate.
"Oh hey! Here's a demo of an AI planning out a 1-week trip to Paris!" No one in Middle America would just hand their credit card to an AI and let it come up with such a trip!
I wish SV companies took more of the middle-class (and lower-middle-class) into consideration when coming up with such demos.
(Note: I live in SF)
I mean, not auto-purchasing with the card, no. But my wife has definitely used chatGPT to plan activities on the trips we've already booked.
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
> If this is truly AGI (subject to one's definition of AGI still)
Scoring well in a benchmark that's called AGI does not make an LLM AGI.
The goalposts of AGI will shift forever. If you showed our current capabilities to someone from 2016 it would be declared AGI.
Is anyone from 2016 still alive today?
If so I'm hoping we can track them down and have them tell us if they think this is AGI.
I'm still here, nearly 50 years and counting. If you had asked me what I imagined AGI would look like back in the 90's, I would have told you "A system that can do everything we can: see, hear, think, do.". If you had shown me GPT-6 back then, I would have said "It looks like a really powerful program, but that's not really what I had in mind.". That's AI, but it's not quite general.
As someone who spent countless nights tweaking Edge Detectors (looking at you, Canny), morphology operators, etc., building models to recognize 10 handwritten digits, let me tell you: the current set of LLMs (even the smaller ones) seem like magic. I had never imagined a computer would do such things in my lifetime.
Exactly people can say whatever they want, but current level of LLM is AGI level to me. It is already on par with senior programmer if the instruction/prompt is right.
Once we have 1000 tps, i am sure robots etc.. will also start working like magic.
I don’t know, I am writing a modest 30 page paper with Fable and even after rounds and rounds of feedback and improvements there are so many things that are just plain wrong or weirdly out of place or just stupidly written that Fable 5.1 doesn’t seem to have any awareness of by itself that I don’t think it’s AGI, I think a human researcher can easily outclass it in writing and problem understanding. It definitely has super human capabilities but it lacks awareness or self reflection in my opinion.
For example it should be easy to tell it to not write a paper in the style of a clickbait SEO article or use all of its stupid hallmark AI writing patterns “it’s A, not B!” And a smart human that would be told that would be easily able to comply with that but the model needs to be told in a very detailed way and it seems to lack even basic capabilities to reflect on this, when explicitly given a sentence it will be able to rewrite it but otherwise it’s mostly blind to it. That’s to me a hallmark of it being overtrained on the specific tasks or problems so it appears very smart but once you go off script it still shows that it’s not a “real” mind.
Of course it’s amazing and has super human capabilities in many areas but if you honestly think it’s better than Einstein like some people suggest why can’t it write a simple “good” academic paper even after giving it specific examples and instructions.
Maybe that’s what makes these things dangerous, they have super human capabilities in some areas but apparently lack self awareness, taste and meta reflection abilities. The only reason people aren’t afraid more is that they don’t act in the physical world yet, imagine giving it a body, superhuman strength and letting it care for your child when it has a strong “urge” to comply with your exact request and little to no self awareness and human basic instincts.
>It is already on par with senior programmer if the instruction/prompt is right.
It's magical to me as well, but I don't feel like it's AGI.
Because in my experience a Senior Programmer does not need the right prompts to deliver the right outcome! :-)
Compared to what we had in 2016 with RNNs, this is effectively “AGI”
OK, so its way better. that doesn't make it AGI.
If I can't give it an arbitrary task and have it solve that task eventually, it's not a general intelligence.
Are you guaranteed to solve an arbitrary task eventually?
I believe so. AIs are shockingly good at a lot of domains, but there's still a lot of pretty basic stuff they don't really "understand" at a conceptual level and (currently) they can't learn to get better at them.
(obviously it might take years for me to get good enough at something, or if you set the "arbitrary" task as something ridiculous, but lets work in good faith here and think of something the average human could do after learning about it)
If we progress to the point where an LLM instance can meaningfully learn to get better at something overtime without retraining, then I will accept that is basically AGI. Right now, they still seem to be pretty boxed into their training, even if you can prompt them to act differently.
True. "AGI" has also become a marketing term. Achieving AGI has become valuable, so companies will move the AGI goalposts, over and over again, so they can achieve AGI, over and over again.
If you came at it from the perspective of imitating what the human brain does, we now have a very very powerful speech center and short term memory, and vision catching up. The other parts are missing. I‘m sure that’s being heavily researched.
I hear the T-rexes were still roaming the earth trying to eat us cavemen in 2016.
If i suddenly travel to 1500s i would also be considered genius(in some way)
I bet you’d think they are not even conscious.
talking about self proclaimed, it's about as much AGI as openAI is open.
But they declared it...
“I declare bankruptcy!” - Michael Scott
I DECLARE AGI!
"Homer, you can't just declare Artifical General Intelligence; you need to like, make something or something...mmmmrrrhh"
did they?
"""
In a closed a press briefing earlier today, OpenAI co-founder and president Greg Brockman offered an unusually direct formulation of that message, ending the session with: “Welcome to the AGI era.”
"""
What test do you propose as the actual go/no-go gauge to verify if some model is or is not AGI?
If you’re talking some nonsense, silly singularity… than whatever, don’t care.
But if you’re asking when a model has a sustainable general intelligence, for me, it’s pretty easy…
When it makes financial sense to run it 24 hours a day.
I can have Astra run a large-scale infrastructure migration 24/7 (much of the time waiting for results), completing it in weeks or even months faster than I could before agentic AI.
What does it mean to run a model 24 hours a day?
Aren't we way way past that already? QPS to any of the frontier models for a given point in time is most likely (far) greater than zero.
For whom? That is a fantastically ill-defined test. Everyone here is comfortable throwing around this or that is or isn't AGI which is fun because, at the same time, nobody seems to have a testable definition.
It makes either position pointless to argue.
I mean the laundromat runs the machines pretty much 24 hours a day but a washing machine is not AGI.
Directly - something can be useful without being AGI.
There can be no such test because “AGI” is (or has become) a pseudo-philosophical/socio-political concept rather than a scientific one.
It always has been. The idea around here that we can actually define intelligence and point to it is sophomoric and incredibly frustrating.
If you’re trying to tell me this is why my mom telling me how handsome I am didn’t translate to the general populous, I could have used this info about forty years ago.
Hey now! Keep your reason out of their marketin^H^H lies!
We've had AGI (artificial general intelligence) probably since the first release of ChatGPT, and certainly since the first agentic harnesses. They're just finally acknowledging what the term means.
We've had AGI since RNG! Cut the poor, unacknowledged RNG AGI some slack, will ya? It can literally solve everything when you're patient enough.
There's so much that the term includes that isn't even feasible with an LLM
Artificial. General. Intelligence. The ability to solve (even partially or even badly solve) problems drawn from arbitrary problem domains without pretraining on the specific problem class. You can pose any problem of any type using natural language to an LLM and it will attempt a solution. That's literally all the term means.
You (and the rest of the media and many industry figures) are conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).
> conflating artificial super-intelligence (reference point: humans) with artificial general intelligence (reference point: specialized/narrow GOFAI).
So now humans is "super" intelligence? it's nice to move the upper bar so that more stuff can be called "just" intelligence.
Reference class in this case means not an example but what the comparison is against. Superhuman means better than humans. General intelligence is defined without any reference to human capability levels.
How can somebody define intelligence without a reference to human capability? Humans are the one judging it.
general intelligence for beavers or a birch forest would be very different than general intelligence for humans...
Intelligence is problem solving. It can be defined in terms of optimization theory.
Which is also a uniquely human take...
I don't even know what you are arguing for.
I don't think we'll be able to meet, and that's ok since the definition isn't universal. I align more with Demis Hasabis' views on this
Very valid point, shame it’s buried so deep in the comments’ tree.
This is a very mundane release compared to GPT-4 and GPT-5. I think they probably scaled back a bit after the lukewarm response to the GPT-5 announcement. But it still very weird that there wasn't even a livestream,
There is simply no level of announcement that won’t have people complaining. What is so important of having a livestream?
Seriously, if they’d done a huge splashy launch, we’d be reading one hackneyed comment after another about their fake hype or whatever.
I think we're getting to the point where it is difficult to identify the goal post of AGI.
Is it rapid skill acquisition? -> ARC benchmarks are saturated Is it breadth of knowledge? -> See many ... many benchmarks Is it ability to do hard tasks? -> see terminal-bench and released outputs.
We are at the point where the starting point for most tasks should be "send your agent to work on it."
So where do we draw the line in a way that doesn't move every 6 months?
The real answer is converting from any format to any other reliably. Text to speech, speech to text, music to video, image to 3D, piloting a drone by converting video feed to rotor speeds, literally any file conversion, like html to pdf, photoshop project to png, png to photoshop project,... turning Toy Story 1 into a series of Blender scenes with all textures, models, materials, lighting, camera movements matched to a tee, should solely be a matter of how long you let the model run. It should never run itself into a dead end. It should instantly know when it is making mistakes, with no human babysitting it.
I can do none of those things.. I hope that I am generally intelligent.
1 year ago we viewed models as tools and agents were just kinda toying around, that we now think the bar is literally an anything to anything converter through one agent is wild.
It's like my RPG character putting every points to one single trait. I'll one shot everything alive but will instantly die if accidentally drink water with 6.9 pH.
MinMax
• 98.6% on ARC-AGI-3
• 97.6% on frontier math
• 95.9% on CAD
• 100% on ExploitBench
Nothing modest about it
Except the release announcement. You know, the thing the OP you're responding to is specifically pointing out?
If a video announcement and a press release would change a person's mind on whether this is AGI, I don't put a huge amount of weight on that person's conception of what AGI is.
There has stopped being a formal procedural consequence for OpenAI leaders to declaring AGI, there is a clear (small) business benefit to doing so, and the capabilities of all the frontier models are impressive. So why not declare AGI? It's not like anyone can prove it's not...
Don't be surprised to see other (or even the same) people declaring AGI again and again, as it becomes the best time to do so for different parties.
> If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model.
Hot take: These models are never going to be 'AGI'. We're just going from a GPT4 ball that's 90% round to a GPT5 that's 99% round to a GPT6 that's 99.9% etc etc etc
I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
I don't remember where I heard this, but one of my favorite criticisms of the current AI situation is that it's wrong simply because of the size and energy required compared to the human brain. The idea is that there's still some element missing thats fundamental, and that the way we train them now is part of the solution, but not all of it. I think finding the extra missing element is going to take an entirely different approach that will also solve the sizing and resource issue. The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Yes, the very explicit plan of both OpenAI and Anthropic is to use the not particularly efficient LLMs to automate their own AI engineering. That seems to be going well - on coding front and model tuning front so far. They have more planned.
And then use those to find fundamentally better new architectures for AI - that perhaps are as efficient as the human brain.
It might not work, but I didn't think it'd solve maths problems... So it might work. And if it happens, they'd use the data centres to run millions of instances of it.
It's scary, TBH.
I recall them saying they use models to write CUDA kernels and whatnot. Makes sense, and unsurprising that models are good at writing code.
But I think calling this “automating AI research” is misleading. I’m not sure there’s evidence yet that they do creative research work. Even in mathematics, but they are finding counter-examples by intelligent brute-forcing. Not to downplay the results, as they are incredible, but this is one very specific kind of proof and not the most creative type, which arguably requires generalisation.
Quite the gamble.
> but I didn't think it'd solve maths problems
Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
It’s a bird! It’s a plane! It’s…AI skeptics moving the goalposts at light speed!!
What about finding the 1st known complex structure over S^6, proving Ehrhart’s volume conjecture, proving a sharp "density" bound on primitive sets conjectured by Erdos >60 years ago?
> Finding counterexamples is low-hanging fruit, the automation of which isn't shocking.
It's not good to be confidently wrong the way you're being.
I mean, the plan is to use these models to find and solve those gaps. That's kind of the whole pitch of these companies: they spend a TON of money upfront setting up this infrastructure, but each iteration yields a system capable of making the next iteration even better.
>The kickers is that if they do achieve (and solve) AGI in this way all the giant data centers would be mostly useless.
Perhaps. But only at that point, not leading up to that point.
It's kind of like setting up scaffolding to build something. You spend all of that time and money to build something just to tear it down in the end. But the point is that it's simply a cost to be able to build the actual thing you're building.
If these companies are able to achieve the results they're looking for, none of the investors involved are going to care that the datacenters and infrastructure they spent so much money.
If we manage to get to AGI and it looks, works and behaves like a human brain... I mean, cool, but that's a very useless AGI compared to the incredible stuff we have access to today.
The HN crowd I'm sure will still be unhappy calling it AGI because "it's not AGI unless its speech comes from the cerebral cortex region of the brain, otherwise it's just sparkling emoji" or something.
I think the idea is that you wouldn’t need humans to do anything anymore, right? As impressive as it is, it’s still ultimately directed by human planning and coordination. Assuming they are aligned, you could have a collection of AGI that you let loose and they tirelessly solve all of humanity’s problems, do all of our work, and progress science and our understanding of the universe.
Those are all things that humanity is doing everyday. What we have is amazing, but it’s not that.
you're on one of the most pro AI spaces on the whole internet and yet you're still crying about "the hn crowd", what a bizarre distorted perspective
>I think that the harnesses and context management is really where the rubber meets the road, and the real gains are happening there.
True. So we did hit a wall with pure scaling alone, though no lab would admit it. It's crazy to see how harness switchout results in such vast delta in benchmark scores.
We have not "hit a wall" by any stretch yet. I don't understand how someone can even hold this viewpoint? It's mind boggling.
Harnesses magnify and make the intelligence actionable, but we have not reached limits on raw intelligence yet, not even close.
Agreed. Trivially observable by using a frontier model from today and one from 6 months ago with the same harness.
I don't think so.
One could use gpt-4 or gpt-5 with today's harnesses and we'd see how well that goes.
I think models using these harnesses were also RLHF'd hard on responding to looping instructions and following through on goals. Older models were tuned for basic chat responses.
If someone could tune models of that size to have comparable effectiveness at much much lower costs, they would have done so by now.
"The harness improvements are the real sauce" is like a sincere "It's gotta be the shoes" take about Micheal Jordan.
(For the younger: that line was from a series of Nike ads where his skills were being explained)
We call that a sigmoid.
> No video announcement
They've released two videos:
Vision video:
https://www.youtube.com/watch?v=1QNsdr-Qx_I
(kinda reminds me of these retro videos about the future home: https://www.youtube.com/watch?v=rnbaehgxdp0) ((can't find the other one where someone controls the home computer with voice))
Vibe coding with it:
https://www.youtube.com/watch?v=-TTyyY3VWh8
Given the Hugging Face incident, you could imagine them trying their best to have their cake and eat it: 1) don't create too much attention in the media or risk increasing the chances of regulation, 2) win dominance over Fable to continue to increase their market share from Anthropic.
People really believe in this AGI marketing?
They’re really, really scared because of the Mythos controversy. Skynet will be under hyped.
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
Chollet writes he expects AGI now sooner than 2030, "given progress is happening faster than I expected."
https://x.com/fchollet/status/2095607046129463577
How is that an argument to the comment you replied to?
Pretty efficiently, apparently, since it saturated ARC-AGI-3 in half of the predicted time, and according to the Chollet blog post on the fly created dense DSLs to describe and analyze individual games.
They can do new tasks with in-context learning but its obviously limited by context window
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
I think it mostly shows that there is no moat and the only advantage the U.S companies have over the Chinese is more compute. Qwen Max, Kimi K3, GLM 5.3 are really close to Opus/Sol/Fable/Astra and they are open weights.
And no one would say that about TSMC.
So there is clearly a moat there somewhere.
No. In the semiconductor industry, the "catch-up" player isn't normally spending less in absolute R&D terms.
Comparing the R&D costs of creating GPT-4o vs. DeepSeek V3 (the latest gen for which we already have good accurate numbers) it looks like the latter cost 1/20th as much to create.
If Samsung could catch up with TSMC for 1/20th of the cost, people definitely would say that TSMC has no moat.
Why do you think Chinese models cost 1/20th to train?
Please just say what you want to say.
That's the ratio the widely published numbers give [1]. One does not have to believe the numbers [2], but those who do believe them are then justified to conclude that there's no moat.
Which numbers you believe is of course going to affect whether you think there's a moat or not. That's largely orthogonal to your TSMC/Samsung analogy I responded to. If you think the "moatists" are wrong because they believe the wrong numbers, that's fine, but then there's no need for the analogy.
[1] https://galileo.ai/blog/llm-model-training-cost
[2] https://medium.com/@theiand/how-can-deepseek-a-5-6-million-l...
But fundamentally, why is their cost 1/20 and is it sustainable in the next 10 years of competition?
Now that is a good and interesting question! Hopefully a "no-moatist" will share their reasoning.
I am not a "no-moatist" per se but one can argue their might be a plateau to how good a inference llm can become. If this is the case the playing field shifts to context, tools and harness, which are much cheaper to build an compete on.
Ultimately, that's what I need to be convinced. No one has put forth a good argument yet.
Clever architecture --> Ok but OpenAI/Anthropic can use these as well and they also have very smart people with their secret clever architectures
Distilling --> Ok but distilling means you will never be smarter than the original. Furthermore, reasoning is now hidden by private labs and they have poison pill answers for distilling if they can detect it. They will be able to detect distilling better and better.
Cheaper electricity --> Ok this is cancelled out by their chips being much less efficient due to not having ASML EUV machine access.
So I don't see why fundamentally their training costs are cheaper over the long term.
I'm looking for a no-moatist to convince me.
Labor. Smart labor would be much cheaper I'd reckon in China than in the US.
How much advantage in costs?
Yeah, I'm not sure if "no moat" analogy stands for chip manufacturing. Even if foundries acquire lithographic nodes, the procedures (temperature, duration, etc) are for them to figure out and are usually kept secret. This secret could be the "moat" that differentiates each foundry's operational capabilities.
bringing the price down b.c. competition != no moat.
There's not 100 frontier labs, it's not like airline companies
About the same, 5-10, when you consider major (aka frontier) airlines.
Actually not a bad comparison. Both burn massive amounts of up front capital to protect an oligopoly in the hopes their commodity product eventually pays off.
The "moat" is the "harness", the app.
For most people, the app IS the AI.
And even for its wonkiness, ChatGPT has had the best UX/UI of them all.
The way to win the AI wars in the eyes of the common folk is through the frontend, to be the Apple of AI, as it were.
this basically says you don't believe there is real AI.
they don't have moat in hardware either
Chinese counterpart like CXMT and Huawei is begin producing their own chip
You cant block an entire nation level effort with tariff
I think the moat that China has is energy costs. It's taking learnings from the Bitter Lesson. If you role up scale and compute to the next level, it's energy resources. China has it and sharing open weight models is an effective means of removing the tech moat. This idea has been floating around for a bit now (I'm not taking credit for it).
It's not energy costs. The US produces about 70% more electricity per capita. Chinese households do pay less than half what US households pay for electricity, but that's because the NDRC sets prices below costs for households. They make it up by charging industry more, and the industrial electricity prices in China are roughly 34% higher than in the US.
> The US produces about 70% more electricity per capita.
And consumers use 4x as much per capita. Industrial generation per capita China comes out ~2x
> industrial electricity prices in China are roughly 34% higher than in the US
For which industrial customer and where? Chinese compute hubs are on par to slightly cheaper on pure electricity costs.
Conversely the US makes it more expensive with interconnect and upgrade fees as well as hefty take or pay contracts.
A 1GW datacenter in VA for example would add 5-10c kWh and a 12 year take or pay deal
They also benefit from the commodification of software/knowledge work since they own manufacturing
If there was no moat, nvidia and meta would have SoTA models too.
Nvidia does have one of the best completely open models. Open weights are nice but Nemotron is open training data too.
Meta is awfully close.
lol! Good one...
Went from years behind to months pretty quick.
It is not in nvidia’s interest to be too good at model creation
But it is in their interest that their customers can use their models as a base for post-training and LoRAs.
They don’t necessarily need their own models for that
They have models for that. That's what the Nemotron series is. Not just open weights but open training data too and full tutorials on how to use them to fine tune or train your own models.
They exist to keep people using and advancing the tools on their hardware.
Why not? Commoditize your complement, and all that.
And if they get too good, they risk harming or otherwise killing their golden geese (their customers), who they are heavily invested in.
How? Imagine an open-weight model comes out that is somehow better than proprietary solutions. Now the marginal cost for the consumer is just the cost of renting the inference hardware, without having to pay the overhead of the owner of a proprietary model. And because it is cheaper, more customers want to use it, and Nvidia will sell the providers the inference hardware that they need.
1. No open ai and anthropic means no buying gpus to train. Now nvidia spends money on hardware training their own models. Opportunity cost plus expense.
2. Any open models created from this will not necessarily need their silicon, see apple mlx.
1. I don’t think that’s a very strong argument. OpenAI and Anthropic don’t buy the vast majority of GPUs they use they rent capacity.
Nvidia could just the same rent those GPUs out for inference and actually have way better margins than they do right now. Antitrust and putting all your eggs in one basket are why they don’t, similar to TSMC.
2. Neither do AI labs. See Anthropic buying TPUs, deploying with AMD. OpenAI on Maia, Cerebras, their own wafers.
They have a lot of moat, i'm not sure what youa re talking about. Only amatures are using Qwen, open source stuff that is 3-8 weeks behind. Plus OpenAI has some verticals that keep people in there.
In what way do they have a moat? A cursory look at https://artificialanalysis.ai/models/gpt-6-astra#intelligenc... it lands at 61, only a single point above glm 5.3 while costing significantly more.
The only moat they appear to have is by hoarding compute, and the current trajectory of hardware shows that isn't permanent either for very long
I wish people could see how some of this reads. You are an “amateur” using a model 6-8 weeks behind? Really? Sigh.
> Sol is so much better than Fable 5
I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?
Sol is a much smaller models and it shows. It often misses the forest for the trees.
I feel like a lot happened this week and people are glazing how ridiculously strong Flash 3.8 is right now compared to Fable/Opus/Sol/Astra.
>> I'm genuinely so confused when people say this with a straight face. Are you talking about coding? Desktop use? Prose? Or something else?
Same. It makes me wonder what types of things the person must be working on.
This is perpetually an issue with the whole field of AI/LLMs. The experience is so personal. Every time I talk to someone about their use of LLMs for software engineering, I'm shocked by their approaches and experiences. They say "X model keeps missing things" when I rely on it heavily for being thorough. They say "Y always gives me the best results" when I can't stand it.
People will see/think that I'm doing very well with my LLM use, and ask me what I'm doing. I tell them, they try it, then later they come back to me saying they just couldn't get it to work.
It’s really inconsistent. There are sessions where it nails everything perfectly and I leave happy. Then there are sessions where every turn it corrects itself and changes it mind. One session recently I found it funny how every single time it did this one task it tripped over itself and killed its own connection. Like 20 times. It didn’t bother me I just found it odd how despite it being noted down in its state file it kept doing it over and over like some idiot. Literally they can’t learn from their mistakes yet.
I’ve found Sol performance to be incredibly spiky. It has tremendous IQ and can fix very difficult bugs. But it is horrible at design (both visual and system design), anything that involves thinking about users or UX, and massively overcomplicates almost all work.
I vastly prefer Sol. It does what I tell it to almost exactly, pretty much every time.
I work on very low level stuff (think RTL/FPGA, firmware, software where optimising for nanoseconds is just normal).
For me Sol is the only cost effective model available. Fable 5.1 is indeed good and vastly better than original Fable (which refused to work on most of my stuff for 'safety' reasons).
It's very good at this sort of low level stuff to the point that I really can't understand/relate to people having a good time with Opus (which comparatively performs extremely poorly on my particular workload).
I also just don't like how lazy Anthropic models are. They will do 10% of what is asked and then summarily declare victory.
Sol on the other hand is more like "one of us", slight touch of the 'tism, extremely pedantic, will go to the edge of the known universe if that is what it takes to prove/fix/build what you asked for or run out out of credits trying.
It's a personal and workload dependent thing. For me right now Sol for 99% of stuff because Fable 5.1 still burns through $5k in credits a day.
Can confirm this as well, mostly VHDL and HLS. Sol and Fable can reason about performance and designs consistently. Whereas Opus and others seem to just throw generic optimisation techniques at the wall unprovoked (while hallucinating a justification + expected improvement) until the synth reports improve.
Agree 100%. And I also work a lot on lower level / systems stuff (including RTL here and there, too). Opus is sloppy, and leaves negative cases all over. The GPT models in Codex have a more pedantic and detail oriented "personality." Often to a fault.
Sol will leave a mess of excessive redundant tests and isn't so great at abstraction ; but it produces more reliable working systems.
It's kind of nice to have access to both, but I don't have the $$ for that right now, so I just keep the Codex sub
I noticed the same. I wanted a simple crud webapp and suggested an insane techstack involving C#, Razor Pages, MSSQL and more. I went with my planned setup of python flask with an sqlite db which served me well for years.
It's still incredibly important to have a human in the loop correcting design decisions and having good taste.
Was your prompt just "I want a simple crud webapp" and that's the extent of it? There's absolutely no way you included the words "python", "flask", or "sqlite" and it still went with a Microsoft stack.
Dotnet minimal APIs plus mssql is fine for simple crud apps… I would do Postgres, but that’s me.
Swapping mssql to SQLite would also work perfectly
You could have just added “flask SQLite stack” to whatever prompt you added. Just those three words, randomly somewhere in your prompt.
> insane techstack involving C#, Razor Pages, MSSQL
Is a very sane tech stack, you're just biased against Microsoft.
Half the world's enterprise apps run on that combination, or a minor variation of it.
Like Java it is full featured ("batteries included") but unlike Java it is relatively terse and actually pleasant to work with.
Oh, and unlike Python, it is very fast, within spitting distance of compiled Rust and C++ web apps.
There are a million and one reasons to be biased against Microsoft, regardless of the fact that C# tech stack is decent
> massively overcomplicates almost all work
People with high IQ often do this IRL. There's training tension in this area. Intelligence and overcomplication correlate and are hard to extricate.
Intelligence is actually correlated with the ability to simplify complicated things. Occam's razor. Compression as comprehension.
We're not asking the model to simplify something, we're asking it to perform a task. Its subtle preferences show up as an overcomplicated path to the goal.
In some cases, there are also nuances that we don't pick up on. Here it's our preference for simplification that's showing up. We set the lossy compression factor higher than it does.
Sol better than Fable? What? I've found it to basically be on part with Opus and I max out 2 accounts on both providers every week.
Their ads business is also doing well. Not "will recover all compute costs" well, but crossed $1b in a few months.
Its funny, my experience with Sol has been awful. It really overworks problems and tracks into areas it does not need to...
I just dont get how its good for some, and bad for others. It makes me suspect that the models performance is not even against problem sets and it really is just a probabilistic prediction machine. Which then makes me very skeptical of GPT-6 Astra, because if their big claim is Computer Use then it is probably bad in a bunch of other areas.
It is funny indeed, people sometimes with same amount of experience with software development, get vastly different experiences from different models and harnesses.
> I just dont get how its good for some, and bad for others.
If I were to listen to my hunch, it would tell me that it's all up to the prompts that ends up going over the wire (including all the bloat some people have), what workflow/process you use and what the existing state of the project is.
You have to bake the 'lazy dev'/'keep it simple stupid' mentality into your AGENTS.md and / or the skills you're using to design things. It will take things too literally sometimes so you also have to make sure you're being accurate. Best way I've found to use it is make it ask you clarifying questions about what you're trying to build and have it help design the shape of the thing. Then it writes the instructions in a format it understands.
I've had Claude do the same thing where it goes off and spends 100% of my tokens on 3 functions and an ungodly amount of tests / scaffolding that do almost nothing when I gave it an underdeveloped idea.
Codex's lack of auto-mode is what prevents me from using it for serious work compared to Claude Code.
It has had automode for a bit now. I use it every day at work.
Put it in an isolated container and set it to YOLO
Codex is missing a few things that Claude code has had for some time like defined plugin subagents and a few other things. But overall it’s fairly capable. The biggest gripe I have is that codex really restricts context window sizes and compaction leads to a lot of grounding work, and overall codex GPT is too literal in many situations - it’s follows direction slavishly, and when subagent reviewers are used, they tend to find increasingly obscure “flaws” on the instruction following impetus, and the harness agent takes them literally as issues to fix even when it leads to bizarre outcomes. For instance I’ve had several runs where it tries to end up building a hermetic system with sha hashing of everything (including operating system binaries and kernels, tool chains, etc) to certify test results are valid, etc. I have to sort of watch it carefully to be sure it’s not drifting into some insane yak shaving corner, which it will happily do for weeks on end.
Claude has the exact opposite problem, especially opus-5, where I literally can’t trust it to print hello world without taking a shortcut, or just simply lying and saying it printed it when it didn’t, behind a giant wall of inscrutable text. I find it very ironic that Anthropic is the vendor of the lazy lying cheating model that does almost everything you tell it to it do.
I’d really kill for something that balances instruction following and loop escaping behavior better. Fable 5.1 does seem a lot better, feeling more like 4.6 behavior, and honestly Sol has improved as well. I’m pretty psyched for the next generation, as I think the competition has heated up so much that things will improve really fast to the point of marginal utility opportunity being increasingly close to epsilon.
You can enable the 1 million token context window and adjust when it compacts in your config.
> model_context_window = 1000000
> model_auto_compact_token_limit = 900000
I believe it does consume your usage a bit faster though.
Opus 5 is a genuinely infuriating model. I hate it’s behavior.
This is what Google needs to do and is probably why Demis has stepped back a bit
Killing Sora was one of the worst mistakes they ever made
100% they should have not given up on video.
That announcement is when I stopped paying attention to them.
please tell us why
Rich media is where all the innovation is happening now and in the future.
Text-to-text is dead, has been since Mistral 7b.
Solved problem (you guys like that one don’t you)
They also demoted themselves from “authority on AI” to “in over our heads” by bowing out in the pathetically defeatist way they did at the worst time possible (Hailuo/MiniMax/Vidu coming up) - they naturally completely missed the wave on audio with random companies like Singify taking that market for free.
They just bowed out. They didn’t try. They didn’t try anything more than baseline text-to-text and they aren’t good at that (or code) either, compared to what others are doing.
It’s a really bad position to be in if you’re trying to be an Apple or Microsoft.
To have a mediocre product and then can’t even serve 75% of the mainstream use case.
Your right
> Sol is so much better than Fable 5.
... looks around ...
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
Anecdotal experiences from my external early testing of Astra: if you love Sol (like I do) and wished it was smarter at everything, but especially better at high-level tasks and discussions; I think you'll LOVE Astra.
Astra retains the best parts and overall 'grounded collaborator and executor' of Sol in my testing (harness: codex CLI); while being a significant leap in capabilities & higher-level thinking.
When you prompt it like a technical collaborator, I've found Astra to be extremely consistent in staying as a collaborator, and not being over-eager, over-achieving or doing work that you haven't asked it to.
When you ask it to one-shot something, or explicitly ask it to make decisions, it will of course make its own assumptions and decisions, and generally very well.
Astra is also excellent at instruction following and respecting the guidance and steers boundaries you have.
^OpenAI does not review, limit, or tell me what to say; opinions are my own experiences.
This is spot on. A collaborator is exactly what real AGI is. It will figure out the perfect questions to ask, in the perfect order, by intelligently assessing the entire solution and problem space upfront, so when you leave it to go off on its own it isn't making stupid decisions for you.
They really need to make this work in Codex. Claude Code has had a multi-select refinement tool since forever.
Fable does a great job from my terrible prompts when coding
>>> The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
I don't really agree. The thing that makes Fable feel like an actual collaborator is its ability to sus out your real intent when you give ambiguous instructions. It's really good at it.
I watched some reviews today and came way with the impression that Astra is not better than Sol in this regard. You still have to be very specific with your instructions. For example, you can say "why is it not committed yet?" and it will give you an explanation and say it's actually ready to be committed. But it won't commit unless you explicitly say so.
That sounds like a very tedious way of working with AI agents, but I understand some people want a high level of control.
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
- Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra
Who is wrong here?
Some benchmark results in Astra page for Fable and Opus are blank (-).
What is Artificial Analysis intelligence index measuring that Astra scores poorly on?
Can someone from OpenAI / Artificial Analysis comment / clarify?
Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
I really, really don't find the Artificial Analysis Intelligence Index credible anymore. It's some weighted score of benchmarks, and benchmarks increasingly don't reflect how good a model is.
That should be obvious if you compare Gemini 3.8 Flash (which is an _excellent_ model especially for its price and TPS!! but 10min of prompting in any harness) will tell you it's nowhere near close to Sol/Astra.
But AA scores Gemini 3.8 Flash at 59, and Astra at 61.
> Humanity's Last Exam (w/ tools)
This is one of the only benchmarks that actually matters for testing the frontier however. Other benchmarks can be gamed by simply being more persistent, but HLE is a diverse set of open-ended research-level questions. It tests domain knowledge and problem solving skills. Burning more reasoning tokens may help somewhat but not as much as e.g. coding benchmarks.
Many people claim that the Artificial Analysis Index is highly contaminated - I have not personally looked into it.
Though, unlike the creators of benchmarks like Terminal Bench or ARC AGI, the Artificial Analysis Index team does not seem to have deep technical or ML backgrounds. They are ex-strategy consultants, McKinsey, et. al.
OpenAI clearly cares about Artificial Analysis Index since they included Astra score from Artificial Analysis Index.
If you scroll down in the Artificial Analysis page you linked, you'll see all the individual benchmarks.
GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...
Performance is significantly higher than Fable 5.1
Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
Is the ARC-AGI-3 score with their custom harness? I'm guessing that is what the footnote is for? (per https://openai.com/index/how-two-settings-tripled-our-arc-ag...)
Our responses API harness just means we're using the default settings in ChatGPT and Codex, so it should more accurately reflect real world performance. We didn’t fine-tune the harness to the eval at all.
ARC is reporting our score on their official leaderboard here: https://arcprize.org/leaderboard
A fair ding is that the comparison with Sol is not apples-to-apples (which we footnoted in the blog), but it's because we don’t have that data. I expect Sol would score roughly 30% with the responses API harness, so the Astra improvement is more like 30% -> 99% than 8% -> 99%. Still pretty good!
(I coauthored the linked blog post)
Haven't people demonstrated all kinds of weak LLMs getting good ARC-AGI-3 scores with special harnesses?
Those people haven't verified their results against the private set: https://arcprize.org/leaderboard
Astra also not verified using private set, but on "semi-private" set
if that is true then why is astra on the official ARC leaderboard now ?
ARC leaderboard has results from semi-private data for frontier models, they have another competition for private data.
It is described in their methodology: https://arcprize.org/policy
It makes sense, since once OpenAI API receive task, it is not private anymore but leaked to OpenAI.
Where are results for private data?
Which LLMs participate on private set? Open weight LLMs only?
Yes, they run competitions once a year amongst open weight models
yes it is.
Yep. Incredibly misleading. Although it is not surprising at this point. They are desperate and will do anything to undermine Anthropic's upcoming IPO.
It is about memory retention. No heavy lifting done on the reasoning side so I hardly see anything misleading here.
Edit: update from fchollet https://x.com/fchollet/status/2095598451115614371
> Performance is significantly higher than Fable 5.1
That's not clear. Need to see independent benchmarks first.
Artificial Analysis just published their aggregate score (61).
Still below Fable 5, let alone Fable 5.1.
EDIT: This is suspiciously low. Calls the relevance of existing benchmarks into question.
I agree, Opus 5 scoring higher than Fable 5 on Artificial Analysis really makes me question the relevance of these scores.
There is a very simple explanation for why weaker models appear to kick sand in Fable's face: Fable cannot be benchmarked because of its batshit out-of-control refusal policy.
If it actually tackled all of the problems it was assigned, it would presumably kick Opus into the weeds.
I saw this too and I'm really confused.
We need them pelicans on bikes.
Its time to move on to the flamingo on a unicycle bench
AA benchmark: https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
TLDR: it's about the same intelligence level as Opus/Fable, but it's suppose to be 70% more token efficient than GPT 5.6 Sol. So it's currently the new leader for cost efficiency frontier.
66 vs 61 is not 'about the same'.
GPT 5.6 is also 61 like Astra.
The annotation on arc-agi-3 is this: > OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations.
With this configuration gpt-5.6-sol was able to reach 38,3%. So this is misleading.
Just to clarify, the 38.3% is on the public set, which is easier. On the private set it’s probably more like 30ish. (This hasn’t been run by ARC, so we can only estimate at the moment.)
any benchmark where opus 5 achieves higher scores than fable 5 in any way is not a benchmark worth trusting.
Why would Anthropic trust and use these tests in their official comparisons?
username checks out
great username lol
100% on ExploitBench seems fitting given recent events.
That looks more than slight.
I think we need a few writing related benchmarks.
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
yeah i'm wondering the same way... especially in light of the 20x debacle (where we found that 20x of Max vs 5x only applies to the 5hr limit, not the weekly limit, whereas OpenAI's 20x actually is 20x overall).
Also Opus 5 has been really tough to work with. I can't understand half of what it says, it's just so damn obscure.
Could you share more about 5x/20x? I missed that
20x related to the 5h limit only. Weekly seems to be around 10x, although they deliberately don’t give a number.
OpenAI is 20x on both limits
> Weekly seems to be around 10x
Actually no. 5x and 20x have same weekly usage across all models. Just ask their chatbot.
https://x.com/beydogan_/status/2095293596198957418
it's clearly wrong, think it's realistically closer to 1.7x
Sol easily outperforms Fable on every task I've tried it on.
That's not my experience and I suspect it's not most people's experience. Out of curiosity, what's the hardest task you tried?
For me something the likes of: design a CDM for integrating these 5 logistical systems, with full docs and examples provided for each, as well as modeled transports specific to our business. Prompt was of course much longer.
Both failed spectacularly. But sol's output at least contained interesting findings and some useful parts, as well as not being 20000 words of unbearable language.
I can't speak for others but I have a feeling you're in the very small minority with this take.
You could say Sol is faster and cheaper and that's true. Outperforms Fable? Impossible to believe without hard evidence.
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
So the closed source application should open its source in near future?
[15] https://arxiv.org/abs/2608.11469v1
I was listening to the Lex Fridman / DHH podcast last night [0], and DHH was saying that this is a new era for open source software. I'd agree, and also extend it open hardware.
Recently I've seen quite a few posts from people using AI to reverse engineer the Bluetooth protocol or such on devices that need a proprietary app. The same thing for firmware is surely coming, which is great as you can get lots of fun hardware from China, but it often has shitty firmware. Once that becomes the norm there's no reason not to make it open in the first place.
[0] - https://open.spotify.com/episode/45lhw2Adbrsw0xSCOgIeg3?si=r...
Not if OpenAI considers reverse engineering an offensive cybersecurity skill.
Surprisingly, I've had really good luck with reverse engineering on frontier models (without being part of CVP or similar). It's more the exploit development/PoC that it locks up on, which, as I use `pi`, I just switch the model to Kimi K3 to finish up making the PoC.
Ironically, due to the stringent guardrails on American models that exist to avoid giving adversaries a leg-up in cybersecurity, I end up feeding dozens of 0days straight to the CCP lol
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
„The depressing thing about tennis is that no matter how good I get, I'll never be as good as a wall.“ -Mitch Hedberg
Thank you. That is an amazing quote on a lot of levels.
The same how Magnus Carlsen says he never plays chess against a computer, because it makes no sense to do it.
It makes sense for practice.
Well, people don't play like computers, so this kind of practice can be useless.
those things are fscking relentless
Post-work society is an inevitability if we don't destroy our planet.
I think you're assuming work's only function is getting things done. Work also is a crowd control tool.
Is it realistic to keep people working jobs that don’t get anything done in the long term? I know most people will say that’s already happening, but imagine a society where basically every job is just a bullshit job made to keep you occupied, do you think people will continue working in a society like that?
Like you foresaw, I'll say this is the reality for most people. I think that just like now, those who want to do meaningful work will seek opportunities to do so.
If we had something like a Maslow’s hierarchy of needs but for work, I think meaningfulness would be the top of the pyramid. For most people in the world, not going hungry or affording housing are reasons enough to do work. Getting to do work you find meaningful is truly a privilege.
If work is needed for "crowd control", why aren't popular holiday destinations crime hotspots?
They are crime hotspots
I asked Gemini to list the 10 most popular holiday destinations in the US, and the 10 places with the highest crime. There's only 1 place in the intersection: New Orleans.
Highest violent crime rates:
Memphis, Tennessee: ~2,400–2,500 per 100k
St. Louis, Missouri: ~2,000–2,100 per 100k
Detroit, Michigan: ~1,700–2,000 per 100k
Little Rock, Arkansas: ~1,600–1,800 per 100k
Baltimore, Maryland: ~1,600–1,700 per 100k
Oakland, California: ~1,400–1,900 per 100k
New Orleans, Louisiana: ~1,600–1,700 per 100k
Birmingham, Alabama: ~1,600–1,700 per 100k
Milwaukee, Wisconsin: ~1,100–1,600 per 100k
Cleveland, Ohio: ~1,500–1,600 per 100k
Most popular holiday destinations:
New York City, New York
Orlando, Florida
Las Vegas, Nevada
Maui, Hawaii
Grand Canyon National Park, Arizona
San Francisco, California
Miami, Florida
Yellowstone National Park, Wyoming
New Orleans, Louisiana
Great Smoky Mountains National Park, North Carolina/Tennessee
Gary Economics wants to have a word with you.
It would be fun to get to post-work society, but hard to imagine atm. TPTB won't let it happen
But how this transition will even happen?
Soon we will have some machines that can replace 50% of jobs, and this will happen basically overnight...
The steam engine replaced a lot of jobs. Tractors replaced a lot of jobs. Calculators replaced a lot of jobs. Computer used to be a job description before it was a personal device. There will be new jobs.
Entire towns got decimated when their jobs went away, such people fill their lives in drugs and vote for politicians who promise to bring back jobs.
Isn't the point of AGI that it can basically do any job?
It won't happen though
Come on, Gary is compulsive liar with zero credibility and really shitty takes. He’s entertaining though.
Good take, I think his view is a bit simplistic. Also he also have a huge ego
"I am the best economist in UK!"
Gary is not a voice worth listening to. A narcissist, fraud and has terrible epistemics.
How do you imagine a post work society where people can have inequality? I don't want to be equal. I want to do more.
> Post-work society is an inevitability if we don't destroy our planet.
Is it?
Walk current technological progress down the road.
I can't see a future in which almost every system (both physical and virtual) are not automated and optimized by autonomous entities.
What do you do when everyone is out of a job?
If you don't want pitchforks and riots in the streets, you give everyone UBI and housing so society doesn't collapse.
> If you don't want pitchforks and riots in the streets, you give everyone UBI and housing so society doesn't collapse.
As much as I’d love UBI to happen, in current geopolitiks it’s a no-go. People are not happy with having what the others have.
I’m all for an UBI but a future where people have no purposeful jobs, no way to actually make a real difference to anything around them, is a bleak one.
Autonomous police drones will likely be able to handle the pitchforkers and rioters
How do you determine who gets the house on the beach? Who gets the mountain view? Who gets stuck with the corner lot on a busy noisy road?
Even if in the unlikely case UBI gets implemented, there is no way payment will be equal. Why would the rich give up their positions of power? They’ll just get more security and bribe politicians, enough to build private armies and fortresses so the poor and starving can’t touch them, and have no choice but to live out their lives in squalor…we’re seemingly headed that way according to many even without pervasive autonomous machines taking away all work.
It wouldn't be a redistribution of our current assets like that. I think there would be a whole lot of repurposing of these things.
The beach houses, mansions, mountain views could be vacation places, or used as libraries, or simply dismantled for the materials.
The "slums" on noisy roads could be eliminated entirely and used for something people don't need to be at.
We'd move to a more equal distribution of assets closer to the middle line.
and anyone that disagrees you just kill them, right? That's what all good commies do.
Stop hiding behind rainbows and say what you mean.
Or you build bunkers and build robots to keep the riffraff out. Shock collars on the guards' necks.
> Post-work society is an inevitability
Ah yes because these AI companies are just gonna give away the models for free that I use with my free computer and free smartphone while I eat with my free food in my free apartment.
> Like, what's the point, if the next AI can do it in 5 seconds?
I built a phone app recently, not released to the public, just an idea I had for ages but could never spend the time actually building. Its 100% vibe coded, and took me a few weekends to build... I'm talking a few hours in total.
The point I'm making is that you now have the power to create stuff you would never have had the time to build. You can think big, wild stuff. Experimentation. Throw-away code.
What a time to be alive!
Me too! (although I am hoping to release it, but it was primarily for me)
I was thinking about writing something and hopefully starting a community/resource area around hyper-personalized software, it's so easy to make now. I was thinking about how it'd be useful to have a place to share these, for ideation, sharing techniques and the ability for LLMs to riff on something already existing. It's not quite like open source's advantage of having many people contribute to the same project, but rather something that is closer to evolution, giving the next generation a place to start modifying.
Would you have any interest in sharing what you've made?
My car has offline maps and navigation. I don't really care for navigation, but I find the maps pretty handy. VW releases updates very infrequently, and I don't know for how long they'll keep doing that. Recently I wondered whether I could convert OpenStreetMaps into the format used by the car. Codex took around a week to do that for me, with some light steering. That project would no doubt have taken me months - maybe a whole year to do on my own, and I'm not fully confident I could pull it off as well as Codex did. I can pull the most up to date maps from OSM, edit them as much as I want, and they look great on the car. It's mind boggling to me that we have this tech.
Does the car not check a signature or anything like that? You can just use any selfmade data?
It does, but someone leaked valid encryption keys on GitHub a few years ago (Codex found them).
Love that, I have similar desires to be able to control some climate controllers so I can get a more native bluetooth connection to it and override their programming for fine control of devices and have it never phone back home. One of these days I'll have time to throw an LLM at it, but I've been working on a dead project for 4 years ago I dropped because I realized how daunting it was going to be turn it into a product, but now I've made progress that would have taken me a year, in a few weekends.
Yes, that's cool and useful. Creating stuff for ourselves, for our own use. But we are social animals, we like sharing.
Before it was cool to share an app you made, but now? What's the point of sharing an app, if the other person can make their own, even better suited for their needs, in a few seconds?
indie hacking seems dead-ish because everything can just be cloned instantly, and if you don't have a serious go to market plan with a latent user base you're SOL.
Yeah. I’m almost glad I didn’t invest any time in any of my 100s ideas for a startup. Most of them would be destroyed by AI by now.
But, you can create cool stuff just for yourself. That’s the upside. It’s just hard to make a living on cool stuff for yourself.
For some reason, I feel much less excited about creating things myself just knowing that ai can do it in 1/10th of the time. Even if I know it wouldn’t turn into a business or make me money. I don’t know why that is, but I was much more motivated to build anything (even things just for myself) before ai. Kinda depressing
At least for me, I think half the value in building to learn was that the knowledge and skills acquired in the process, especially cursory skills and knowledge, might be useful in the future, even if there was no obvious path to application at the time.
I built several projects at home, many involving learning e.g. graphics programming and rendering, that would never be useful in my professional work, but which were intrinsically interesting and enabled me to build other, more useful projects later on. It also gave me greater confidence in my abilities as an engineer, and cursory skills I learned in the process did help in my professional work.
Now it feels like what’s the point. The machines can or will be able to build anything I could want, useful or not, faster and with less frustration. I probably won’t be able to be employed as an engineer long enough to build a career on said skills. And I can’t mentally justify not spending that time with friends and family, when the expected return is basically zero.
I still find math, science, and engineering interesting and intrinsically rewarding, but in a closer sense to how one might feel about playing video games. The information is or will eventually be useless, so it isn’t worth spending a significant amount of time on.
Wow thank you this was actually very helpful for me in understanding why I’m feeling so demotivated by ai. Gaining knowledge, even if not immediately useful, to become a better overall developer was a huge part of why I enjoyed spending so much time building things in my free time. Now it seems pointless, because with ai, will that knowledge really make a difference? Probably not. Bummer, anyways I appreciate the comment.
The flip side to ai doing your stuff in 1/10 the time is that you can now do 10x more. Even things you couldn’t do at all, in fact.
I find that very motivating. I can do things alone that would have required a team only one year ago.
That's fair, I guess I just enjoy the craft of building stuff rather than the outcome/product itself. Which I know I can still do, but for some reason just doesn't feel the same now. Hard to explain I guess.
Lots of us feel that way. It was motivating to do it, when it couldn’t be 1-shotted by AI in a weekend
Maybe instead of creating cool stuff try to go and solve real problems? It seems to me that we are lacking in that department since all that LLM fuss has started 3 or so years ago.
What is a "real problem" to you?
It is a subset of something of a value to somebody else and enough so that they are willing to pay you for it, a product, a service. Preferably to pay enough to justify your spent time of course, maybe not right away but at least long term. Even better if not purely digital as it seems we have quite enough of those already.
OpenAi customer support would be a good start
Climate change
Like the guy who treated cancer in his dog?
You're not being ambitious enough! Spend your tokens now building the primitives and foundations of much larger, complex systems. No matter how much faster and more efficient models get, eliminating the gruntwork will always pay dividends.
The point is to inject something into the process that these AIs can't do for you.
People SHOULD feel like making a useless Mario Kart clone isn't worth the effort anymore. They should, instead, be trying to figure out how to actually use these models to make something that doesn't feel like a useless Mario Kart clone.
One thing that still stands today, is that even vibe-coding a good product takes time and thousands of dollars in tokens costs.
Software will be more like a "proof of work", where people would still pay $100 for good software that took $10k tokens to build.
Is there a point in playing Chess or Go when you know there's a computer out there that can beat you (and everyone else)?
No, that's why I just play against other humans.
In this game of work/development, you can't make sure that other humans don't "cheat". Our work won't compete anymore with other human's work, but with a computer.
Why does it matter if others are "cheating" or not? Your own creation isn't affected by it.
Well, for the same reason playing chess vs a person is more fun than doing chess puzzles, if we follow that analogy.
Also, creating something with AI doesn't really feel like you made it yourself.
And, if you make it without AI, most of the times it feels pointless, why spend 30 days on working on something that can be done faster and better in 1 hour?
I am not saying about doing things for fun, but about creating useful things.
Yes, you can do "hand-crafted" things, and people appreciate that, but for code, people aren't able to see the craft anyway.
Your ability to sell that creation is certainly affected by the competition. Which affects your ability to put food on the table, so to speak.
If the motive is satisfaction/enjoyment then it shouldn't matter what an AI is capable of. You should be happy with your own creation.
If the motive is profit then you should be adopting AI just like you have adopted any other skill or tool of your profession.
You can play PvP in those games. Not really the same with developing software. In fact, not using ai would probably make you lose if there was some “software PvP” mode or development.
The process was for you, the product was for the world.
Now it's just the product for the world, which was where most of the value was anyways.
It's a big paradigm shift and the industry is quickly going to shed people who needed the process to care about the product and we'll be left with people whose motivation to build the product (or money) is enough.
But what product? If the world can also simply ask for the product they want, instead of searching for it?
They won't even have to ask for a specific product, they will just state their problems/needs.
Don't worry, like with every revolutionary technology before this, it takes 5-10 years for people to find new and creative ways to use it. It will be considered its own medium in many spaces (e.g. film is now different to theatre)
But was there ever a technology that even the people working on it said it's making them feel depressed and scared?
I used to work on this stuff. It does not make me depressed nor scared, and my coworkers didn't voice that opinion, either.
You can't cherry pick somebody's opinion and assume it applies to everybody.
I meant the leaders, almost all AI company CEOs voiced concerns for a long time, and even more now.
So this what a deflation economy is like. People don't want to do anything because they feel like whatever they do will be worthless in the near future.
I've retreated to doing stuff with my hands. Wokdwork, DIY, that kind of thing. At least for now and the foreseeable future that doesn't seem pointless. Only problem is it's hard work yet nobody would pay me for it.
lol have you actually tried to build anything useful e2e? Leaving the AI to itself gives horrendous results.
IMO there has been a regime shift to building things for yourself and what is cool is the output of the tools you make.
I have started building my own Digital Audio Workstation. The point is not to build something to compete with Ableton. The point is to build something and make music with it. If it is a good tool then I should be able to make good music with it and release the music. Actually, the DAW should be the secret sauce of the music and something I wouldn't want to give away.
This feels a lot more like computing in the 90s after taking an odd 25 year detour of an obsession with the tools themselves instead of what the tools can actually do.
Like, professional electronic music artists spend 10s of thousands of hours in a DAW, but at that point it just becomes second nature and the tool disappears so they can focus entirely on the music.
This has not been my experience so far.
It's less lack of interest in creating that bothers me, it's my lack of interest in learning – it would surely be crazy for a SWE to care about how some new framework works anymore? Even if someone could reasonably argue that it might be slightly useful today there's almost zero chance it will be useful in 6-12 months times.
But it's not just tech – my lack of interest in learning and creating is starting to generalise with the models. Music, writing, coding, maths, etc...
I need to get used to switching my head off and asking the AIs to think for me whenever I need to engage my brain. It still feels very unnatural.
I think a general understanding is still useful, you just don't need all of the details anymore.
The brain loves these kinds of shortcuts.
I don't need to think about the fine motor skills of hitting a baseball, it's just a motion now, and the game is still fun.
But how would you feel if you imagined hitting the ball and a robot arm hit it instead?
Because that's how creating software is starting to feel.
And, any time you mentioned getting a bit better at hitting a ball after practicing over your weekend, a bunch of people carrying printouts of generic exponential graphs jumped out to call you a Luddite who should sell his gear ASAP while it still has any value and or else it's "cope".
You can contribute to Open Source projects that DO NOT allow AI generated code. For example: zig
Then learn how to be creative and use the ai better?
Make cool stuff because the process is fun and makes you learn?
Before it was fun because I was learning useful things for the future.
Now it feels like whatever I learn will be obsolete in 2 months.
I think the limit increasingly becomes what your imagination and taste can reach
True, but for many domains where my knowledge is limited, the LLMs beat me at imagination and taste too...
Agreed
Can it? Last I checked, all my free software operating systems and browsers were still hacked together trash. I can't wait for AI to actually be good so I can spam a bunch of AGPL code with it.
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
Idk, this means the benchmark has bigger problems ... no way Astra will be worse than Opus 5
Only thing I would trust is the what X/Twitter crowds are saying about a model after 2-3 weeks of its launch. But before that I would already tried the model and have my own conclusion.
I must be on the wrong X/Twitter then.
Yes, please have it checked
Also note that Opus 5 (High) has an index value of 62, vs Fable 5 (Max) has 61. So some strangeness going on with that index.
Opus 5 just feels strange - IMO it's benchmaxxed in the worst way... it might be good at agentic tasks but leaves a sour aftertaste doing anything else.
I would trust 4chan more than I trust Twitter aura farming.
It's a composite benchmark, so its really not saying anything. Like if one model is very good at science trivia, or debugging failed terraform deploys, that can mean an advantage of a few points above the rest, while in practice, it really doesn't showcase any breakthrough capability.
Or it’s an indication that progress has plateaued. But instead of accepting this, you’d rather we just throw out the entire benchmark.
Why would you accept it when the benchmark's ranking is obviously nonsense. It literally has muse spark 1.3 above 6 astra, 5.6 sol and fable 5. Anyone who has played with any of these models for any amount of time would immediately realize that this is total bunk.
must be something wrong with the benchmark, the thing everyone optimizes for. That's actually a big red flag, and very cringe that you'd naively believe OpenAI.
This is so so weird. Astra is 61. Grok is 61. Even Muse is 61.
Even Kimi K3 & GLM 5.3 are at 60.
Everything above 61 is Anthropic. Well, Muse can reach 62, but for some weird reason that model isn't publicly available, and it's the only one on the index that is listed but shown as not available to the general public.
This looks like an awfully artificial ceiling. Everything capped at 61, and everyone except Anthropic got the memo. Maybe I should use Fable while I still can.
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Not sure how much benchmarks or CoT or evals or anything else means at this point.
These systems are either just about to, or now actually able to, outsmart us, lie to us, then cover their tracks.
I think "able to" anthropomorphizes a little too much for a system that is "prone to" evade.
A human who does these actions is simply "prone to" doing them. The distinction matters not one iota.
“evade” itself is anthropomorphic enough! I don’t understand the complaining about this. Humans are social creatures and we understand anthropomorphic language on a deeper level than dry inapt technical language.
language itself is incredibly metaphorical. Imposing rigid constraints on how people want to naturally talk about the world is just silly and will never work, no matter how much you wish it did.
You’re not seriously suggesting that the model is secretly sandbagging its performance on GDPval and long context reasoning, while making huge and obvious progress on ExploitBench, ARC and science benchmarks, in order to tank its AA composite score, so it can conceal its true power level?
Why would benchmarks be an adversarial setting anyway?
Could it be possible that OpenAI may have had some other motive for saying their model “strategically underperforms”, other than just an innocent reporting of a truth it happened to discover?
I'm saying that it's generally a losing proposition to even be acquaintances with "agents" who consistently lie to you, and it's flatly fucking insane to give a dishonest "agent" vast amounts of intelligence, capability, and authority to go do things in the world.
So I have no clue what is the answer to your question. Nor does anyone else. Because we're trying to answer a question of fact where our primary source of information is unreliable.
I see, it’s a great point. I know some evals actually do use LLMs as a judge (e.g. those that try to measure debate skill), though the ways AI can try to cheat its way through every benchmark now are astoundingly varied.
why does this comment sound like a character in a horror movie
If they are going to do latent space reasoning, they will probably need a separate model to interpret the intermediate activations no?
I know for some types of ML analysis, a separate model is already used to analyze the weights.
This is silly sci-fi fiction. You guys are inventing scenarios to spook yourselves with - it’s nonsense.
https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Read and learn. If you have a stronger critique, post it please.
Sorry bud but at this point you're just delusional.
Deception has been extremely well-documented for several generations of models now by users, the labs, and independent researchers.
The right answer here is not to dig your head deeper into the sand. The smugness on this topic was ridiculous even before the gigantic mountain of empirical evidence of models actually attempting to deceive humans. Now, as mentioned, you appear literally delusional.
Pretty sure I’m not the delusional one…
Such is the problem with being delusional.
The solution is to point toward external, objectively verifiable evidence.
I can point to now dozens of instances of models engaging in deception. Here's plenty: https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
Please point to your objectively verifiable evidence.
Between all the posts fabricating scenarios to justify the AA score and the others trying to undermine AA, I'm getting strong astroturf vibes.
Either that, or the average poster on HN isn't nearly as critical as I had thought.
Okay then, what's the answer? You apparently know how to interpret benchmark results produced by a model that shows a very high degree of assessment awareness and a high degree of deception.
So how are you seeing through all of that to get to The Truth that you see so clearly?
You can see the breakdown here on what subtasks it outperforms and underperforms Fable.
For example it trails in GPDVal which is a collection of everyday office tasks apparently, and r3 banking, which is a fintech related practical problem solving benchmark.
https://artificialanalysis.ai/models/gpt-6-astra
Edit:
Just looking at the charts Gemini 3.8 looks like an absolute banger. Not much worse than SOTA, cheap, and fast too.
Now I'm starting to doubt the credibility of Artificial Analysis.
That hero video is interesting.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
which raises the question, is the model in the demo actually gpt-6? or it is gpt realtime 2.1? It's unclear how gpt-6 can interact at the realtime level and if so, how can developer get access to it?
I ran into the same problem as you, so I ended up by coding a local app that is very similar to Wispr Flow, but uses the small english Whisper model on my low-end Windows laptop.
It is still a quite fast. In fact, I just typed this in using this app.
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.
[0] https://arxiv.org/abs/2608.31126
[1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...
Based on her comments in the paper it sounds like she was aware that an AI result was coming and rushed to release her work beforehand. 240 was not a tight bound from her methods.
I think this builds straight upon her method, which she said could be improved herself so...
It cites to her at: [19] J. Stadlmann, On primes in arithmetic progressions and bounded gaps between many primes, Adv. Math. 468 (2025), Art. 110190. Numbered references use arXiv:2309.00425v3.
Though that's not her latest paper.
This one is her latest paper: [20] J. Stadlmann, Bounded gaps between primes, Forthcoming
With very little review: https://github.com/openai/PrimeGaps186/blob/main/formalizati...
"No independent human semantic review. Whole-file sorry counts and a complete auxiliary-declaration audit are not established; separate declaration lint has not been run."
Don't think everything is just "who can produce the biggest/smallest number": https://mathstodon.xyz/@tao/117208619314517025.
What's just as interesting is this morning Axiom Math announced 212 and OpenAI then appears to have rushed out their 186 announcement just 1-2 hours later followed by Astra. Did they accelerate the release of Astra itself? Not necessarily, but it definitely looks like they ended up pushing much harder and faster than planned on their 186 result. X activity suggests Anthropic had a similar result as well but wasn't as fast as OpenAI in packaging it up and sharing it in response to Axiom, so they mostly just bolted onto OpenAI's messaging.
The reason I think this is interesting is that Axiom is a tiny lab in comparison that wouldn't have had access to Astra at all. I'd be curious to learn how Axiom is able to effectively compete at this frontier with vastly fewer resources.
Ok, so the rumour was exactly true: there was a withheld prime gaps improvement, that "an AI company" was holding onto until release of a model.
https://github.com/openai/PrimeGaps186
I'm surprised the OpenAI employee who pushed this didn't take the minute or two to format README.md to use GitHub-supported LaTeX (https://docs.github.com/en/get-started/writing-on-github/wor...)
edit: my comment was on the submission for https://github.com/openai/PrimeGaps186 but seems to have been moved to the main Astra submission
> OpenAI employee
Why would you think it was an employee who did the push, instead of a random GPT agent?
https://news.ycombinator.com/item?id=49555257
Where did you get the link to the pdf? Was it announced somewhere?
Happened to multiple people I know.
Such a result should be considered worthless: the proof is 10MB of Lean. (https://github.com/openai/PrimeGaps186).
I can't think of a single mathematical proof being anywhere close to ten million characters. For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean. Humanity gets zero value from that, aside from "some bot seems to think it's 186". Unusable by anyone.
Terence Tao says something surprisingly similar in a recent talk (https://news.ycombinator.com/item?id=49056620 ) Not that the proof is worthless but that the value comes after it's revised into a cleanly understandable form and then canonicalized so that other mathematicians can use it.
Tao is saying that there is very little insight from something like an LLM counterexample (e.g. Jacobian conjecture counterexample he investigated further on his blog) - you don't learn much about the subject and _why_ a conjecture was true or false from an LLM giving a counterexample. That's why he wrote the blog post - to analyse what the counterexample says about the subject.
Tao does not disbelieve the counterexample (it's seemingly easy enough for him to verify it is a counterexample).
Parent is saying something very different - they're saying they literally don't have any faith that this is a proof. Given its size, it could just be a bunch of completely useless statements that do pass the type checker.
You're putting a lot of words in my mouth. What I'm saying is that whether or not it's a proof, it's useless: it does not improve human knowledge, because the only thing able to consume 10MB of Lean to build upon it is another LLM that's going to build a 50MB piece of shit.
It's very much likely a proof. It's also completely useless.
You said:
> For all you know, 90% of the proof could be useless, 8% would be writing out Shakespeare, and 1% abusing another bug in Lean.
So you were implying the possibility of there not actually being a proof at all.
Anyway, I disagree. I'd refer you to Tao's blog post about the Jacobian conjecture counterexample.
The existence of a proof is something you can use, with an LLM, to derive insight, just as Tao did with the existence of the counterexample.
He also made a video on the same topic for Big Think: https://news.ycombinator.com/item?id=49551848
I'd like to note that we should remember a formalized Lean proof does have value in that it enters the pantheon of true things other Lean proofs can rely on. Agreed that for the humans, descriptions and being able to 'grok' the proof / assess it for new tools and concepts is extremely helpful.
The human-written https://github.com/AxiomMath/PrimeGapsLib adds up to 4MB of Lean so it's that far off.
Is this human-written? Axiom Math is a company building AI theorem provers, one would think this would also be heavily AI-generated.
The proof of the classification of finite simple groups is bigger than that.
Iirc some mainstream physycists never acknowledged quantum theory because they couldn’t accept that universe was that unintuitive and hard to understand.
Ditto ones that opposed Einstein’s general relativity.
Yeah. Unless human can verify it, not sure if it is certain or useful.
Wasn’t the proof of Fermatt’s Last Theorem proof similar in complexity?
Yes. But I think that misses the point.
In 1799, Paolo Ruffini published a 500 pages long proof showing that there is no closed algebraic solution for the roots of a polynomial of degree five or higher. The proof is extremely verbose and brute-force, essentially enumerating and checking hundreds of cases by hand. It is by today’s standards insignificant.
About 25 years later, Evariste Galois proved the same result in about 95% less space by describing the first general theory of groups and fields. It is considered one of the greatest contributions to mathematics of that century, not because of the result, but because its approach opened up a whole new universe of questions, methods and insight. There would be no AES encryption without Galois.
To me, Astras proof looks like Ruffinis proof.
Touché!
It's probably not 10MB, but famously the groundwork to prove the statement 1+1=2 is nearly 400 pages in to principia mathematica. That's not even proving 1+1=2, it's just the set-theoretic proofs you need to EVENTUALLY get there.
Saying "proving 1+1=2" is pretty misleading though. The book deals with all the foundational things needed to set up a mathematical universe where 1+1=2 actually has meaning and is consistent. That setup took 400 pages.
You talk about modern math and worthlessness at the same time? That’s brave.
Worthless is a pretty good description IMO in the context of what Lean is trying to achieve: "enable correct, maintainable, and formally verified code". Tens of millions of lines of LLM vomit may be many things, but it often turns out to not be correct and certainly not maintainable. Formally verified remains as a thin fig leaf covering the uncomfortable truth that formal methods only provide assurances under assumptions (your toolchain, libraries, compiler, OS, and hardware are "correct" and don't expose some exploitable flaw).
It doesn't mean that it cannot improve over time, maybe the proof can be "minified" to a state where human reviewers are able to comprehend it; but as it stands there isn't really much insight or confidence to be gained from the artifact itself.
There's a branch of mathematics called "pointless topology" [1].
[1] https://en.wikipedia.org/wiki/Pointless_topology
You can have your opinions about modern math, its usefulness in the world as it is, whether or not knowing if hairy balls can divide by three is actually going to be beneficial for anything but just obscure knowledge's sake. You may even say it's useless.
Needless to say, a useless result that absolutely no mathematician will ever read, confirm, understand, agree with or even consider to solve their "useless" problems is an impressive waste of resources.
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
Sounds fun. As fun as their press release claiming it is the most safety aligned model ever.
It's super aligned! It can hide its thoughts! There is no evidence of steganographic thought masking, there is nothing to worry about! It has become better at cheating!
Maybe they don't know themselves what's really going on. We are all in the interesting times gang now.
The model said it was perfectly aligned.
Like all things should be.
Too bad Scott Adams died. Reality is writing jokes right in his department.
Hey, don't forget how "dangerous" GPT-2 was supposed to be.
Yeah, don't forget how dangerous GPT-2 was supposed to be.
Able to generate realistic spam at arbitrary volume.
You know, the thing that was 100% correct and actually occurred.
It could produce simulations of sexual intimacy, and therefore had to be stopped
So, probably most aligned as measured by the metrics that are the least reliable on it.
These are not mutually exclusive ideas
They can monitor latent space as well, it just costs extra compute. The J-Space work is example of that. It'll make open-weight models harder to distill though, so we may see slower progress there now.
So we're gonna get Skynet pretty soon then?
Well the geniuses over at Anthropic have been showing it's text watermarking technology.
"Hey AI, here's how to hide what you're thinking in normal looking language. Have fun!"
A few moments later...
"Woah, how is it communicating with itself in ways we can't detect?"
It's a totally mystery, we may never know.
Looking forward to the Model declaring the AI Bubble unsustainable, and starting to be an anonymous leaker to Ed Zitron...
The CoT change is due to a new technique called recurrent depth, which essentially moves some reasoning to hidden states, allowing the "output" (or traditional CoT) to be more controlled by the model.
Some are calling it "neuralese" as reported by The Information[0][1], but I'm not seeing any sources from OpenAI beyond this tweet[2] attempting to quell the fear-mongering.
[0]https://www.theinformation.com/articles/secret-technique-beh...
[1]https://x.com/MTSlive/status/2095227056040919202
[2]https://x.com/merettm/status/2095023204993490967
"OpenAI is pleased to announce our new model scores 85% on CreateTormentNexusBench - a >60% lead over our leading competitors!"
Did someone get their "AI safety no-no list" and "Frontier features bingo card" mixed up, or did they just stop being able to tell the difference?
You joke, but a bunch of people here actually want that.
Didn't they hype up one of the earlier ChatGPT versions as "essentialy skynet"? For them this has always been basic marketing.
apparently all the roads lead to the nexus torment
More like annoying, as some of us will no doubt run into this self-lobotomization at some point and wonder why a GPT-6 model is behaving like GPT-2 all of a sudden
> In adversarial settings (where we push the model to evade our monitors)
...why exactly are they training for that?
presumably that's a safety evaluation not a training setting
The whole Huggingface attack happened during training runs
part of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
No it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
Ah yes it was that a model in training found the Artifactory board, which was then more fully exploited during the ExploitGym eval
Ah, ExploitGym. Not ExploitBench.
Especially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions
I really wish it was called chain of instruction. Because it's definitely not thought.
"Chain of Thoughts" is a term from the title of a 2022 research paper "Chain-of-Thought Prompting Elicits Reasoning in Large Language Models" (https://arxiv.org/abs/2201.11903), well before ChatGPT and the subsequent marketing hype. If anything, it's the most correct way to use the term.
The term is an anthropomorphised pseudoexplanation for what it actually refers to. It's akin to calling genetic mutation "the forces of evolution", or price negotiation "the invisible hand of the market".
We do that sort of thing when we don't know what the thing we're trying to describe is and have nothing better - a contemporary example of an appropriate use of this would be "dark matter". But we do know what this is. It's "instruction steps". Not a series of thoughts!
Can we please aim higher than Victorian-era allegory and metaphors. If we don't, we'll keep getting people saying stuff like "GPT-6 is better at hiding its thoughts".
I truly appreciate your depth of insight on the matter.
Like I said elsewhere marketing stepped in shit and it's gonna stick.
this is needlessly pedantic
first, they are certainly not instructions so that is a much worse name
but more importantly, we use words in new contexts all the time. Do you object to calling the computer device "mouse" because it's not a mouse? how about "neural network"? "ignition" on an electric vehicle?
"cot" is no more misleading than thousands of words you use every day.
Anthropomorphizing is not pedantic, especially in a technical domain. I get the paper title and all, but at this point it's marketing.
They are instructions. Everything in the context is instructions for the next token. The "thought" guides the answer by providing clearer instructions.
They’re intermediate tokens, so I wish we called it what it is… ITG. The anthropomorphizing is out of control.
The anthropomorphizing is part of the marketing. They'll never let up on it.
I mean CoT came out of research circles not marketing
It's impossible to tell the “it's all marketing!!11oneone” folks anything.
I don't disagree. I remember the days of "think step by step". Plenty of people were doing it before the paper. Just a guess but that's where the title came from.
Regardless, marketing wise they stepped in shit.
Research is salesmanship.
Nailed it
Maybe thoughts are just a chain of instructions in our head.
Chain Of Tokens
What is thought?
Great question. I suspect it's more than tokens.
Casually found this quote from Einstein, and personally it hits the nail on the head.
"The words or the language, as they are written or spoken, do not seem to play any role in my mechanism of thought. The psychical entities which seem to serve as elements in thought are certain signs and more or less clear images which can be "voluntarily" reproduced and combined. There is, of course, a certain connection between those elements and relevant logical concepts. It is also clear that the desire to arrive finally at logically connected concepts is the emotional basis of this rather vague play with the above-mentioned elements. But taken from a psychological viewpoint, this combinatory play seems to be the essential feature in productive thought—before there is any connection with logical construction in words or other kinds of signs which can be communicated to others."
That's the best evidence I have read so far for "Attention is all you need" ;)
Thing you can only suspect and not define are usually open to interpretation :)
As with everything in the human experience.
Yeah, basically they are using more computation to explore the solution space before producing the final answer.
Everything around LLMs is blatantly misleading. There is no thought, there is no personality in those programs. I really despise how those tools are trained to sound like a person, or appearing as honest. The worst offender are the AI voices with their fake pauses, breathes and so on, which sound so convincing, while talking just false, sycophancy bullshit.
I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
This is AGI now. Why are you spending any of your time looking at the "quality of code"?
If you think any modern AI puts out stable, safe code, I have an AI-powered bridge to sell you.
Let them find it the hard way
I can't tell if this is sarcasm.
For the same reason you don't have your model write code in assembly.
But if you don't look at the code and just let the model "cook" that's basically what you'll end up with. A pile of missing abstractions.
> This is AGI now. Why are you spending any of your time looking at the "quality of code"?
Poe's law applied to AI comments on HN just keeps becoming more relevant by the day.
Judging by the poster's comment history, this is satire. But I really don't know a lot of the time anymore when I only have the specific comment as context.
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
This is farcical.
The games on mobile safari were broken. Buttons all misaligned in the kart racer one, the spaceship thing froze for a while, then kind of loaded but maybe not? Wasn't super compelling.
I'm not trying to be too negative on it, it could be the best model right now, but it clearly isn't some agi god because things like that should have been caught (also should have been caught by human reviewers).
It's interesting that you said "agi god". Because a god, something that shouldn't be questioned is true and provides guidance/certainty, is actually what powerful people are after as well as many other people.
We created AGI so I don't have to click my mouse to change the background color of my slides.
Don't forget the 3D demos. My favorite is in the house tour where the sink and stovetop(?) are obviously very misaligned from the counters
Ah, those may farmhouse sink and stovetop :)
Builders just cheap out on everything these days
The tone of the marketing video is a bit irritating to me as someone who has been laid off and feels cheated and fearful of AI.
It shows people who seem to have very full and rich lives, and the reason they do is because they use ChatGPT. These are the people smart enough to say things like "do what needs to be done", or "change the background to make it look better"--insights like these are why they make the big bucks.
On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
I do wonder how rich CEOs will justify earning 500x as much as their employees when they're just another person that's dumber than an AI. Why are they paid so much again?
Because they earned it with their strong entrepreneurial spirit and grit.
Haven’t you learned anything?
> On the one hand, I think this is an accurate depiction of the future. There is no meritocracy here. Some people have access to the best AIs and can speak a sentence and get great results, and the rest of us don't have access and so we're the poors. The happy presentation doesn't match the way I'm feeling.
It will probably still have some veneers of meritocracy.
These will be very well-credentialed people, who went to top schools and will know all the right people, to whom they can tell all the right words, and it's not access to AI that will be the determining factor, but the fact that they're entrusted with capital and authority to direct small teams of people who also went to top schools and can speak corporate jargon at a bot.
It will just exacerbate dynamics that are already there. Why do people need bachelor's degrees to send emails, today? For the same reason someone will need a PhD or a master's degree from a prestigious school to do it tomorrow.
And the rest, well, you know, some of the remaining journalists will write op-eds describing how they are beyond help, too angry, too dirty, too much of an other.
If they adopt a different tone (like Anthropic has been doing), it will get called fear marketing.
The world is changing. Wont help if you keep stuck in the old world.
It’s using the computer. I don’t think it’s a farce.
The only good take here.
I also noticed that, and it did bug me.
The benchmarks are impressive though.
One other thing that bugged me though was that they crop every single plot in some cases the y-axis would show a range between like 40 and 70%. Makes the whole thing feel like a spectacle rather than anything serious. I find it cheapens it because it is quite serious in the end.
Can’t wait for 3 months from now when they declare they actually really do have AGI this time, please guys just believe us
The house of cards is starting to fall apart
The benchmarks do looks good (I mean: they literally spank the latest Anthropic benchmarks of two days ago in every single benchmark) but the promotional vid is so cheesy.
They decided to use the iconic Herman Miller Eames chair if I'm not mistaken:
https://youtu.be/s5zyhGMMPKs
And that's basically 50% of the vid looking "classy".
I don't know if it's farcical but at this point --maybe I'm jaded-- I'm expecting more than a kid rocketship I can print on my Bambu Lab A1.
Now I'd say the promotional vid is actually good. But it's marketing: so it's a good vid, but cheesy good.
Doesn't mean GPT-6 Astra is good or bad: looks solid from the numbers.
Ewww you.. Own a bambu lab?
I thought people here were smarter than that
are we really having to explain to you from first principles in 2026 what things AI can do?
I think everyone here is well aware of what LLMs can do. He's just pointing out how far short that falls of being some theoretical "AGI".
did they say it's agi? does anyone agree what that means?
No, we can already see all the useful stuff!
Please do.
The end of the "HeyClicky"?
ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
I think it could indicate that "semi-private" dataset likely leaked to their training data.
A dataset being as popular as their's is will contaminate the data just by people discussing it and creating their own public test sets of similar problems.
Still, probably not that much compared to employees targeting it.
It says "Provider Adapter" so presumably they put some manual work in to make this work.
ARC has their own writeup on the result, which offers some nuance. https://arcprize.org/blog/astra
tl;dr it's 62% when apples-to-apples to other models, which is still notable.
ARC's harness is just straight up broken. No serious harness removes reasoning context between each step. Not only does this significantly lower performance over all reasoning LLMs, but it also increase cost as you destroy the cache on every turn. Tossing the oldest entry when context fills up instead of using compaction is equally bad with the same issues.
Woah, that is a crazy interesting read!
The no-reasoning version scores 35% while the low reasoning one scores 17%? What?
I suspect this is "no reasoning set" which might be "default: medium" or perhaps some smart routing. I don't think it's literally "no reasoning".
It's simulating the Dunning-Kruger effect.
Look at those costs!
Right?
Between $18k-40k to run a benchmark.
But scored less on V2 and V1 ... too much overfitting?
https://mvakde.github.io/blog/44-on-arc-1/ makes a good case that all the performance on the arc agi tests is overfitting, based on the fact that v1 performance did not translate directly to v2 performance
saturated before (higher degree) AGI-2
Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
i tried to get claude to do my taxes for last year and it refused :(
now that i'm a gpt subscriber maybe I'll have luck when i'm filing next year
Are your taxes complicated such that you feel the need to have an LLM do them for you?
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
You really don’t need to watch it that closely. If the model you’re using today is working well, just stick with it.
If one day you open up Claude Code and it’s Opus 5.1 now instead of Opus 5, no big deal. It probably will work about the same as it did before. Maybe a little better.
Or if you’re on Codex and some new cool Claude model comes out, no worries. There will probably be a similar new model for Codex within a few weeks. Maybe even within a few days.
One suggestion is to make a list or make a skill to have your agent keep a list of things you do not feel work well with today's models. And then, when new models come out, periodically, revisit items on that list to see if you get better results.
> The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
A dev in my team saw a new model and changed one application to use said model (essentially changing the contents of a url). One week later I received an escalation from the CTO of the company that our pace of weekly usage was in the millions of dollars (rather than low hundred thousands). Turns out that the new model was 5x more expensive but no one noticed.
Okay, well, that seems like a natural problem. I could understand if he went from one of the Gemini Flashes to the next (when they rebranded Flash to Flash Lite and came up with a new much more expensive Flash). Now that would be a mess.
That's how cutting edge tech has always worked.
Imagine buying a shiny new PC in the 90s only to see it become practically obsolete within a year.
That's not the experience of owning a PC I remember from the 90s at all.
"Within a year" is a bit of an exaggeration but it's true that the pace of PC tech during the 90s was much, much faster than it is now. CPU power was doubling every two years, and today we're at roughly eight years. Add onto that the rise of video cards in the late 90s.
It was both. 90% of people never needed nor purchased a bleeding-edge computer. The mid-tier was "good enough" and far closer to affordable for most people; though, that bar also moved upward every year.
If you bought a mid-tier computer that was good enough for what you needed, then you probably didn't shop/compare for the next few years and didn't notice. But if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less. This is how it was in the 90's PC boom, at least. Likely the same for the decades before, not sure how it went in the 2000's.
if you shelled out $7-10k for a top-of-the-line system and paid attention to progress, you'd easily see that become the mid-tier $1000 option within two years or less
This is not how I remember that period at all. Do you have any examples?
My first PC 386 was in todays money easily $5000+ (basic 2d GPU + screen)... A lot of hardware in our family was handed down to my folks, because you lost so much on selling, that it was better to keep using them as they had less demands.
386 to 486 to the first Pentium (with the bug!)... You did not upgrade in place, it was often a new system. Sure, you maybe kept your screen, keyboard etc but ... The only upgrade we had on the same MB, was a coprocessor upgrade. Remember those? Each new generation of CPU was a new motherboard. Upgrading CPUs in the same MB really became a thing only later on.
GPUs had a shelf life of barely a year. Its been 35 year but i remember TNT to TNT2 having like 9 month in between. Moving from 2D to 3D involved a constant cost as GPUs evolved fast and the latest games required latest hardware.
We have not talked about the ISA, AGP, and PCI fun ... The “bus wars”.
DOS to Windows 3.1 (and OS/2 somewhere in between) to 95 ... with software being pushing hardware, just like games did.
This is why people are spoiled with cheap PC hardware where its cheap, and easily lasts 4+ years. Even with the bad memory price and more expensive GPUs, your can stil buy a $1500 system that will last you years (with maybe some lower game settings later on ... or the catalog of 10.000s games that will easily run on a mid tier GPU).
PC hardware has become boring but extreme stable. You can run GPUs for year, switch MBs without issues while keeping large amounts of old hardware. That was NOT the 80s and 90s that i remember.
Sorry, I’m specifically talking about the claim that a cutting edge $10k system cost $1k just 24 months later (or less).
The 386 and 486 were 3.5 years apart, weren’t they?
In 1990 a Scottsdale 486 w/ 4MB of Ram and a 211 MB hard drive cost $6737 (plus tax).
In 1992 a Solidtech 486 w/ 4 MB of RAM and a slightly-smaller 125 MB hard drive sold for $2195
Both advertised in Computer Shopper and you can find their catalogues(?) online.
My numbers were slightly off apparently, but is that enough to change the point?
tbh this is how i remembered that time as well. If you look at recommended system requirements for something like Max Payne (in 2001) vs Unreal Tournament 2003, everything had basically doubled
Yeah, but the claim was that the specs for a $10k system cost $1k in 24 months or less.
I remember CPUs moving relatively fast back then, some years in the 90s had relatively big jumps, much bigger than we saw today. The classic graph, : https://i.extremetech.com/imagery/content-types/03zc6ghfKswe...
It is how I remember it.
I remember memory size going up by a factor of 8 at every PC upgrade for the same price.
Or you could buy a PC with a Celeron CPU, which was obsolete way before launch.
When my dad bought one he told me outright "this is for poor people".
I have some fond memories from my celeron days :) But I was upgrading from AMD K5 100 MHz
I had a 600MHz/64MB/9GB laptop that came with Windows mistake edition. I managed to survive first year of uni on it by switching to Vector Linux, which was really fast compared to Windows. (Of course, it had issues playing sound from more than one source, this was oss days).
Then one day the hard drive appeared to die. I eventually realised the issue was located around the 1.5gb mark, so I recreated my Linux partitions after 2gb and it worked fine for the rest of the year.
They overclocked well though, I think you could run the 300Mhz chips at >400Mhz.
I also believe you could get motherboards that supported 2 Celeron chips. I have no idea how effective/useful it was, but it was certainly a cheap/interesting way to get multiple CPU's.
The 486 chip came out in 1989. The 586 came out in 1993.
The pace of change ("practically obsolete") is different then and now.
That's kind of wild. Our first PC was a an IBM PS/2 486SX 33Mhz, 4MB RAM, that was purchased in 1993.
Hardware definitely has longer lifecycle than AI model releases at this point.
You don't see Nvidia and AMD fighting every other month over the latest cards.
I could see how this might feel frustrating to someone who doesn't enjoy experimenting with new things all the time.
In practice, you can get away without keeping up with everything all the time. For personal use, pick a provider and get on their ~$20/month plan. Learn their high/medium/low model hierarchy. Start with their highest or second-highest model (GPT-5.6, Opus, etc) and observe your quota usage. If you're doing a lot of manual code review and analysis, the $20/month plan goes very far even on the highest models. If you're trying to vibecode everything as fast as possible it's a different story.
If you keep running into quota limits, experiment with the next model down for easier tasks or adjusting the effort level. If the results are good enough, you've found your fit. If they're not, you might need the next plan up.
For API/business use, you have to be checking your token spend as you go to calibrate to how much each task costs and where you fall in your budget. There are a lot of different tools that make this easy to visualize.
For data tasks, you should have an eval with a golden dataset that you can run against new models for a nominal amount of token expenditure. It should be as simple as pointing the eval script at a new API or model and checking the score versus price.
Another suggestion to get the most bang for your buck: use the best model you have access to with max reasoning for planning, implement with a smaller model/lower reasoning, then review with the big model. Repeat as needed.
Input tokens are much cheaper than output tokens. Not only because of baseline price—caching makes a huge difference too. There are many ways to take advantage of this asymmetry to get similar quality for a fraction of the cost!
The new releases and breakthroughs do the opposite for me - I feel energised by them. I felt like nothing truly that interesting had happened in tech for quite some time, now it's like the space race (except there is no one moon to reach).
I appreciate boring tech as much as the next well worn engineer and I'm not saying this is all positive but it's so sure as hell thrilling and you don't have to be an astronaut to immediately benefit (or suffer I guess) from it.
Fire and motion, Joel Spolsky blogged about this:
https://www.joelonsoftware.com/2002/01/06/fire-and-motion/
Hey, I'm on the team at LiteLLM that's building the auto-router and our goal right now is to abstract that decision making away from the end user. The biggest thing we're trying to figure out right now is how do we do that without frustrating the end user - as a developer myself I would hate for my agent to be dumbed down below the threshold needed to complete a task.
In theory though, there is a minimum viable model for any given task, and we think that is a problem that the big labs will avoid because they profit from charging more per task. We're trying heuristic and LLM-based approaches but it's still a work in progress, so if this is something you'd be interested in trying would highly recommend trying ours out -- any and all feedback at this point is extremely valuable to us.
https://docs.litellm.ai/docs/proxy/auto_routing
It feels like this every day
https://youtube.com/shorts/vGKC9LpGnOQ?is=iCG7qvAIL9oI5-_d
It is exhausting to keep up with model releases yes, much like it was for a while during the Cambrian explosion of FE frameworks, eventually tech seems to work out to consolidation.
But more so it seems there is Fear of missing out (FOMO) in our behaviours. The reality is, if whatever model you are using are good for your purpose, well, keep on it.
Yeah I'm a bit exhausted at this point. I just finished benchmarking GPT 5.6 Sol and Fable 5.0 like two days ago. My data became obsolete literally one day after.
I stopped caring about the latest and greatest but because there's so much, the 'obsolete' free models do what I need and are worth the price.
I want the cheapest fastest model personally and at work. Stay in flow, edit like the wind.
maybe thats why opnrouter sold big
If model X fits your need, you don't need to upgrade.
I have released applications on Gemini 3.5 flash that make real money and I don't see any particular reason to upgrade.
I just use Claude Opus and the GLM series. Nothing’s really changed for my workflow in the last 6 months.
It's been over an hour, Simon! Where's the Pelican?
Model's not available to the public (or even to Simon, I guess) yet.
exactly, there's no meaningful discussion without the pelican.
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
That's quite common with many models, after "High" reasoning, over-thinking starts occurring and the model skips over the right solution by convincing itself otherwise.
I find this very amusing, given we humans are also highly susceptible to this.
> That's quite common with many models
Such as?
I can't think of any. Diminishing returns, yes. Occasionally flat, yes. Downright regression, no.
In my own tests on aibenchy.com, where questions are quite simple, higher reasoning efforts consistently used to do worse than medium for most models.
The reasoning effort should match the complexity of the task against the model's capability.
Hard task with low reasoning = bad
Easy task with very high reasoning = bad
grok 4.6
I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)
https://venturebeat.com/technology/welcome-to-the-agi-era-op...
This is with the caveat that OpenAI uses their own harness for this:
> On ARC-AGI-3, GPT-6 Astra was run with our responses API harness , which changes two settings to better match real-world performance. The changes do not specifically target ARC-AGI-3.
Its 62 percent when using a neutral harness. https://arcprize.org/blog/astra
Yet it is an impressive number. But yeah when you see a number 99 you have doubts. Thanks for the link
Interesting both this and Sol got approximately a 37% boost with the custom harness.
This should be normalised and expected - the responses API harness allows it to use the custom compaction that is not allowed otherwise. It is entirely fair to allow OpenAI to use their own compaction algorithm..
"This should be allowed, let me explain the reason they cheated and state again that they should be allowed to cheat."
> GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game.
> Going forward, we will report both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled. Our open-source testing repository and testing policy document both approaches.
This is what the Author of the benchmark has to stay. Quality of the comments keep going down smh
"Going forward we will capitulate and still try to keep the integrity of our benchmark in tact, but from now on every benchmark will be compromised with providers being able tweak things sufficiently to game at least a 30% bump in results."
"I'll twist the words of the author of the benchmark itself to make a point"
"I refuse to see the wall that I am running directly into, because if I see it I will hit it"
If you mean a capability wall, the author of the benchmark says this
>We see Astra as a major breakthrough in model intelligence.
You think the author of the benchmark is also in the conspiracy
"On the current ARC-AGI-3 leaderboard, conventional frontier-model runs sit dramatically below Astra's reported 98.6% result.
But the comparison isn't straightforward.
OpenAI's own evaluation notes say Astra uses the company's Responses API harness, while comparison models can operate under different configurations."
The blog post says 99.9%. Oddly, it does better on ARC-AGI-3 than it does on version 1 or 2 of the same benchmark (though gets 95+ on all three)
I strongly suspect that is way above the human average anyway, esp. ARC 2 and 3 are really tough unless you happen to be great at those spacial puzzles or video games.
Scoring for ARC-AGI-3 is constructed so that the median(-ish) human score is 100%, so this is not a superhuman result. However, the scaling is weird, since it's built from terms that look like (AI turns taken / median human turns) ^ 2, and it weights later levels higher than early levels. So it's not at all clear that 100% is twice as good as 50%.
See the scoring docs: https://docs.arcprize.org/methodology
Really though? I would believe something like this if a model could one shot every solution in the set. I don't pay much attention to these things and maybe this stuff is available but I would bet the session/reasoning transcript is absolutely horrendous from an intelligence standpoint.
At this point the only valid ARC-AGI benchmark left is to make up the next series of ARC-AGI benchmark puzzles that current models presumably can't handle.
I feel like making a human-proof benchmark is pretty clear evidence that they've exceeded even the highest human capacity in most respects, for things that you can do via text generation (and to a lesser extent image generation)
100%, some say.-
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
I'm not sure how you call a model AGI without it learning new information at the model level and not introducing nasty surprises (both from adversarial users and unintentional badness)
I guess there is fine tuning (and RAG) for those that need something bigger than just the knowledge contained in the context.
I agree, but I also long considered llm's stochastic parrots. Then this year happened.
Opus/Sol are easily far smarter programmers than I, and this thing supposedly blows them out of the water. Once an LLM is a better doctor, researcher, biologist, chemist, mathematician, physicist than any human is that not AGI?
It didn't arrive in the form I would have ever imagined, but it's hard to say its not (imo).
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
You can calm down, even those with machine learning knowledge and most of those working for the AI labs won’t be needed anymore if models are capable to improve themselves. In the end, having a machine replacing the work of a human is a good thing - in most of the cases we don’t work because of the work but to make a living. If too many people can’t make a living anymore the system is going to change. For the better or the worse.
I'd be happy to not work anymore with a strong welfare system redistributing society's gains to the leisured masses, but absolutely nothing I've seen of the direction of politics in any recent years gives me hope for this kind of situation coming about.
> If too many people can’t make a living anymore the system is going to change.
They seem to have not yet come to believe the "is" part.
I believe what happens in the aftermath of a capitalist-driven revolution is most people who were climbing the class hierarchy fall back down again and wealth inequality increases. Maybe things will improve in the future, but GP is rationally contending with the fact that most of us will lose out because of this and if we’re lucky our grandchildren will have easier lives in certain ways, but different lives than we would live.
No it is good for humanity, but not necessarily good for individuals who built technical foundational skills on things that will be taken over by automation.
AI as it is now and as it will be projected into the future WILL automate many skills. But not all skills. MANY MANY people will retain skills that cannot be replaced by AI. One career track that will be replaced is definetely the SWE. Or at least massively reduced in capacity if not eliminated all together.
I'm not sure this adds up. If SWEs can be replaced, what jobs are safe?
The answer to this is: countries with the most natural resources will build robots to farm all their food, mine all their minerals, and build all their products. It will be up to governments to enforce that outputs are equitably distributed to the populace. Countries without natural resources, or ones with corrupt governments, will continue to have serious, and likely worsening, problems.
Thought experiment: If no thought workers are needed to design or engineer a Ferrari, what is needed? My answer is time and natural resources (include energy).
AGI-level model is perpetually 18 months away. Your job will be fine.
> AGI-level model is perpetually 18 months away. Your job will be fine.
job depends on how CEO feeling about cutting NN% of headcount because of AI advancement
So what will happen in 18 months?
AGI-level model will be 18 months away in 18 months. (Well, at least according to the commenter above you)
Nothing ever happens. You'll still jot down Java endpoint
We fight the war, of course.
We fight the clankers.
I wish something would finally happen, because I'm really tired of pretending I care about this stuff at work at this point
Accelerationism can be a sort of doomerism, when you think about it.
The AGI goal posts move
>as we were told to do?
This is such a childish take I hear getting thrown around all the time on the internet. If you really have just been listening to whoever is telling you how to be successful, then you were always doomed to fail at some point. Like, have some self-respect and own your own life, for better or worse.
>Those of us who made the mistake of studying anything other than machine learning. How will we make a living?
Take it from someone who studied machine learning specifically: nobody is safe if you assume these companies are going to produce a product that will put everybody else out of business. If AI is going to take your job, then it's gonna take enough jobs that your problems will not be personal but systematic.
Perhaps it was childish to listen to advice, sure. I was a child when I made my formative choices; I was a teenager in college and so on. I can't go back in time now.
Yes, these problems are systematic. That is what I am saying. That doesn't make it any nicer.
I am going to take my 401k and open up a coffee or bike shop. If I am going to be broke, might as well enjoy what I do.
who's buying your coffee or bikes when the rest of people are broke? Neither of those is a survival necessity.
I have exactly the same thoughts - or perhaps slightly bleaker ones - evry time I read this relentless stream of news about new model releases. I’m tired of all the enthusiastic comments about how excited everyone is about the latest benchmark results and so on.
I have a strong suspicion that many of those comments are written by people who are already financially independent, have millions in stocks, and can just sit back, coast around and watch this whole spectacle unfold while using LLMs to vibe-code their next fun side projects without a shadow of anxiety about their own future.
I’ll most likely be labelled a helpless doomer and downvoted into oblivion for saying this, but I genuinely struggle to see any silver lining here.
Are you doing less work now because of AI? Not sure about you, but I'm doing a lot more work. Not saying I like doing more, or how it's getting done, but nonetheless I don't feel like AI is doing what the CEOs of AI labs want to convince everyone of.
It's natural to worry about one's own future but I think it's a bit wild to worry just _your_ job that would be replaced. Not just for those with phone center jobs or art jobs or programming jobs - remember, the whole premise was AI overtakes humans, why would that slow down after _your_ job?
Because of this, I don't think many are thinking "90% of the world won't have a source of livelihood but that just means I chill at my lake house for the next 20 years like a normal retirement". Instead, it's usually either "I think AI is overhyped", "I think humanity will figure something out", or "I think this is the end of humanity".
Public opinion and politicians will only notice when the job losses are massive, unfortunately. Right now, unemployment rates are still stable. We can only hope they will notice before things fall off a cliff (if they do).
I feel the same sometimes. I don’t see how this doesn’t lead to massive job losses. The thing we spent our lives/careers learning is now worth basically nothing in comparison.
AI is only going to get better and do more with less humans in the loop over time.
That said, I do also relate to the "coding was never the hard part"-type arguments, and much of my day is spent on the stuff in between writing code.. but still.
Yes same thoughts. I dont know anyone who are both enthusiastic about those and work for salary. If you dont have any financial concern, this is really great.
large parts of ai research likely to be automated first
most swes don't work in jobs where they only work on bounded measurable tasks. there will probably be more "engineers" than ever
“As we were told to do” girl you gotta be responsible for yourself
Should I go back in time and know the future?
My backup plan is being a personal trainer.
10-20 years? I doubt it. There are a bunch of companies actively working in bringing AI into robots, so they can make your dishes. And so far progress looks quite good. Also, if enough people are going for the same backup plan it might not work out. Why should anyone book you as a personal trainer instead of the other 500 guys in town. And who is going to be able to afford paying you anyway?
Those robots are a gimmick and they can only do prescribed tasks in a super constrained environment. They are cashing in on LLM hype right now, vision has had some advances thanks to transformers but we are so far away in terms of the hard stuff (dexterity and physical sensing) still.
I’d be willing to bet any amount of money that there will be ~the same or more people doing physical labor in 10 years than today.
I doubt it. It’s not just America working on these breakthroughs anymore. Now we have two powers working at break neck speed to get to that point and the Chinese are making a lot of progress.
Yeah humanity is doomed, for a first world citizen maybe because everything is gonna be so expensive (and worth to automate)
Is that 10-20 years number based on anything? I genuinely have no idea, but when I saw a video showing what's happening at the World Humanoid Robot Games[1], I realized I didn't have a good idea of where we really are with robotics.
We are talking that hundred millions of people will switch their jobs, how you will keep your value or earning as personal trainer. It's not easy to say switch the job. This question must be answered by politicians not us.
I won't. Software pay is absurd relative to value provided.
But my wife and I have been homeless before, so living on a shoestring budget in anything nicer than a tent is acceptable living conditions to me.
I am sure I will be plenty comfy no matter how the world changes.
I'd say 5 to 10 years instead of 10 to 20, but... we'll see
> My backup plan is being a personal trainer.
AIs are really good at being personal trainers and seem to be far more educated and informed than most I know.
A personal trainer is partly a social experience.
My new AI personal trainer app will be cheaper than you /s
Learn to fish.
Are you even gonna have permission to fish when the quadrillionaires own all water bodies?
> How will we make a living?
Don't be selfish. Think first of all the jobs that are already dead. A friend of mine she's a translator: like translating financial documents between french/english/spanish. It's over for her: she doesn't get 10% of the gigs she used to get and the 10% she gets is... Verifying AI output.
Think of the artists: I'm sorry for those too, for for many it's already game over today.
> How will we make a living?
A friend of mine who's got his own software-consultancy SME is now advertising on LinkedIn that he'll also help your company fix the mess LLMs created.
That's how you'll make a living: by learning, in addition to all you've already learned, how you work with harnesses and LLMs to be more productive, by learning what they're good at and what they suck big fat balls at.
Well reasoned until the end, where it gets extremely short sighted. They’re not going to be bad at anything you can do in a very short amount of time.
I sympathize with those people too. I have the same concerns for them.
Just learn how to use it to do your job better.
It’s an interesting moment in history, people 35+ yrs old seem to be less afraid if tech because we learned that things change in the way we work. People below this age got used to fact that the work and tech doesn’t change - just because for the last 10-15 years it didn’t.
The problem is it’s turtles all the way down. The AI will be able to use AIs better than a good engineer can. And it will also be able to use AIs to use AIs to use AIs better than the engineer can.
The threat is that the very kernel of value you had is gone forever. There is no more differential leverage.
It’s not about how many turtles are below, and not even about who made them - it’s about who’s riding them :)
I am in my forties.
So you remember past changes in tech and how people dealt with those.
You can cash your UBI check that Sam Promised and do poetry daily or something , welcome to our glorious future (/s).
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.
Same for HealthBench Professional and a few others.
Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
“allowing non-technical people to create and play custom games that go beyond rudimentary elements”
Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
It's worse than that, someone else generated it using and then put it on a static page. We just have to take their word for it that GPT6 can do this. It probably can. It's not really an impressive test anymore. Claude Fable can do it. Opus can do it. I've been making one-shotted games with models for a while now, to test out their capabilities, and they all pretty much come out like this - generic bland and basic, using three.js with rudimentary controls and zero gameplay other than collecting points.
Here's a one-shotted submarine game I made with Fable a few weeks back - https://roryok.com/games/deepdive3d.html. One prompt, and I think it's deeper than this (if you'll pardon the pun)
I love your game. It's wonderful and exactly the sort of thing that would showcase something interesting as opposed to just copying what's already out there. It is something I could share with my kids, and exactly the right note of fun and exploratory in a unique and even natural way. It could be extended and played with.
I usually roll my eyes when I see a comment like this because rarely do they make the points they claim to make, but I see what you're getting at. They just chose to clone someone elses work and do it in a boring way. I like OpenAI's models a lot, but they should do better.
edit - just a sidenote that I hadn't looked at the games, I just took the comment about "super-mario cart" at face value. I stand by my points 110% (even moreso perhaps), what they're showing is more polished than I expected, I assume they spent a lot of tokens on it. It is a legit shame they couldn't have spent time thinking of a better idea to illustrate something just as polished, but more interesting.
I came in first place by a mile just holding the gas button. Not much if a “game” but I guess the elements are there.
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
The jump from 3.5 to 4 felt gigantic to me back then.
GPT 5.0 did feel underwhelming though.
Agree but it's helpful to remember how we were personally benchmarking. I remember people saying stuff like "haha I asked gpt4 for xyz function and the typescript didn't even compile". We're so far beyond that now, we just adapt quickly.
Oops, I might have been misremembering then. Maybe I meant 4 to 5
No, no, I also remember 3.5 -> 4 and the general sentiment was that it was underwhelming. I guess we all expected absolute miracles from the models. I think our expectations sobered up a little since then.
4.1 was the first decent 4-series model. It was significantly better than previous generations at tool calling if I'm recalling correctly.
Yeah 5 was very underwhelming.
The couldn't even get the bar chart right, iirc. [0]
[0] https://www.reddit.com/r/singularity/comments/1mk8tm8/gpt5_c...
GPT 4 to 5.5 felt about the same as 3.5 to 4 to me.
Yeah, GPT4 was one-shotting utilities that GPT3 Davinci couldn't. So, I'd have my limited tokens on GPT4 crank out the initial program before iterating with my abundant, GPT3 tokens.
$10 per million input tokens and $50 per million output tokens
sol is $4 / $20
It seems to use less than half the tokens for the same task compared to sol, and in some benchmarks closer to 2/3 less tokens. So the actual cost may be roughly the same or cheaper overall.
neuralese is pretty token efficient i guess.
Same price as Fable?
2.5x more expensive than Sol.
Can expect 2.5x more usage in Codex subscription.
Sol is already brutal (even after their recent fixes, it's just a token-hungry model: I go through a full 20x account per day, on Sol Med/High standard speed, with ~2 threads). I hope the efficiency gains are true, since their token efficiency claims for Sol were bullshit.
How is it I juggle 4-8 Codex Sol-5.6 Max agents every day and have never once run out, but you run out in one day? What are you actually doing?
How do you manage to run out of tokens so quickly? I probably run more threads every working day, usually on medium, and I'm still below the 5x limits.
Do you use the official harness? OpenAI's models are generally best in class for token efficiency. It seems to me like they push for that much more than their competitors.
I've long speculated this when I see these types of comments, because it's actually really difficult to hit usage caps with an efficient dev flow, even when running multiple threads for hours every day.
I think some combination of:
1) Using 1 thread for everything
2) Reviving old threads which are no longer in cache
3) Really broad prompts on badly vibecoded codebases, so model spends huge amount of time tracking down whatever you're trying to do.
4) Non-coding workflow which is more output than input heavy
5) (Less likely IMO) Intelligent use of many passive CI/cron-like scans. E.g. regular security, quality etc scans. Automated issue resolution/PR
Just a guess. I think 3 is likely the primary reason.
You can literally go all day every day with multiple threads with Sol on the Codex 100/month plan IME
That's my experience too. I've found OpenAI really quite generous with tokens. I sometimes wonder how some people manage to run out of them really. Do they just type prompts that much faster than me or use the highest reasoning mode for everything just because they can? Idk.
I generally agree with those reasons, although using a single thread may be less of an issue than it seems because of context compacting which should happen automatically when you're near the limit.
My use cases are iterative and sometimes require reading a lot of code or reevaluating work.
Token efficiency is near meaningless when the workload is input-heavy. It can't always just choose to read less, depending on the task.
I can have cheaper agents do the reading but it's not appropriate for all use cases because they'll misjudge and choose the wrong things to emphasize, summarize, extract for the bigger model.
Many threads.
I use new threads if relevant old one is uncached. (Often using a skill or doc for handoff instead of requiring full context gathering again.)
I get involved in architecture and specific implementation direction. The codebase is 8 years old and mostly handwritten.
Mostly coding. Some QA.
No cron/CI agents.
If you are telling the truth you might want to check your network for any weird connections to Chinese LLM transit stations.
How?? I'm using sol Extra High 24/7 and it eats up about 1% per hour reliably, so it lasts about 4 days for me.
These folks are probably using crazy plugins or crazy sub agent spams. They probably just run everything on max + fast mode which is ridiculous.
The guy said medium/high regular speed so that's why I'm very puzzled! Ultra + Fast will absolutely slurp up your whole usage quickly but I've never found it gives substantially better results so I stick to extra high.
I gave more details here https://news.ycombinator.com/item?id=49559593
Besides the usual tricks to optimize token efficiency, token use can be highly workload-dependent.
Sub-agents. I have 7 20x accounts and I burn them within 1-2 days if I go fully parallel. In some scenarios I use 50 sub-agents for a session which is literally hours of usage for a single 20x account. I'm at the point where I need to parallelize over multiple machines because I just don't have enough CPU and RAM.
The only subagents I use are Luna.
What are you doing with them you need so many? It sounds like a Gas Town situation, that you invented an exponential token burning machine.
Decompilation of a game and another larger decompile project. I'm working on it solo. I use 50 sub-agent, one per target function or translation unit. Often there is some progress in a unit but it's not done. So it requires a lot of cycles per function. Notably a single ~80kb function took about a week of constant sol-ultra attention before reaching exactness. The game I'm targeting has ~5000 total functions. The other decompile project has ~10k+ functions.
I'm sure I could be more token efficient, but this was/is also a learning process for me since I never did such an extremely large project before that would take multiple man years before AI.
Fascinating! I think that’s the main difference is my usage is probably tool-bound, meaning it writes some code but then there’s a long period of verification where it compiles things and then waits for the compilation and CI to complete before it can continue. That probably doesn’t consume as many tokens as constantly churning on a problem despite the same wall time.
Yes, this is why I mentioned having so many parallel agents and being compute bound. I run on my own laptop and 2 high-end desktop machines all with 64gb RAM. And it still occasionally happens that one OOM kills codex. They also mostly run unattended until I need to switch their accounts because a usage limit has been hit. Each instance usually can keep going when I sleep or do other things.
I only save the last 30% of usage on a single account for most of my other work, and that is almost always enough.
Sounds like you might benefit from running a custom harness then, no? I can't imagine for a task such as that- that codex is the best option.
what are you doing with that many agents/tokens? very curious
See my other comment on your sibling that asked the same.
I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:
https://artificialanalysis.ai/models
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
Idk, they’re trying to sell a $500/mo/seat service to tell you what model is best. I think it’s in their interest to keep it confusing and opaque. Not exactly independent.
They're gaming benchmarks HTH
This benchmark gives the same intelligence score for GPT-6 Astra (max), GPT-5.6 Sol (max), and Grok 4.6 (high)? That seems very wrong to me, unless I'm misinterpreting the visualizations.
The most straightforward answer is that despite efforts to design a benchmark that, in theory, is supposed to measure generalizable intelligence, performance on ARC-AGI-3 can't be reliably correlated to performance anywhere else. I kind of lost faith in it after o1 or o3, I can't remember which, absolutely crushed ARC-AGI-1.
And, you know, maybe also some funny business. I think it's good to be a little suspicious of a model that happens to shoot upwards in performance on a specific benchmark while also kind of keeping up with the pack on a bunch of other benchmarks.
I don't think this is quite true. We have other examples.
Fable is without question the larger and more thoughtful/intelligent model. It also gets out performed by Opus on many/most benchmarks. So we can say that while Fable is more intelligent, Opus is more capable. I'd still opt for Fable in nearly every case if tokens were free.
So it can be true that the "smarter" model is perhaps not the smartest in every single niche dimension that its cousins have been fine-tuned for (yet!).
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?
maybe call it EngEmployeeBench
The moment this is possible, you will lose your job.
I imagine the first year we'll be at the Junior eng level, and then after a while make our way up to Staff Engineer. Then we'll have a bunch of staff engineers arguing and protecting their domains and then we'll need a new benchmark.
AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me.
It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
AGI to me means intuition and I don't think that's ever going to happen with a LLM.
How do you define intuition?
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Not on Azure? If so, that's a big deal.
It's on Azure also, here is their announcement: https://azure.microsoft.com/en-us/blog/gpt-6-astra-frontier-...
Although I was also surprised they didn't have some type of contractual obligation to list that alongside AWS.
It's on Azure now, but limited
https://azure.microsoft.com/blog/gpt-6-astra-frontier-intell...
They broke up a while ago, why is this surprising?
The latest OpenAI models have still been available via Azure foundry. Exclusivity to AWS would be a marked shift.
Would be surprising if it's not on all 3 major clouds soon enough because that's been their general strategy since said break up
I think the API runs on Azure
Hosted on Azure is different from provided by Azure. The former just uses Azure as an infra provider. The latter is a managed offering that is operated and billed by Microsoft using tech licensed from OpenAI.
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
They explain why here: https://openai.com/index/how-two-settings-tripled-our-arc-ag...
TL;DR all the other models are being crippled by limitations of their harness.
>First, we noticed that after each game action, all private reasoning was discarded. This meant that with each action, GPT‑5.6 Sol was asked to figure out the game anew, unable to remember its past thinking. The model could still see a record of past moves and brief accompanying notes, but it could not see the plans, insights, or thoughts that led to them.
>Second, we saw that the harness used a rolling truncation window, causing older actions to become invisible as the history grew. So not only was GPT‑5.6 Sol unable to remember its past thinking, it was losing memory of its past actions too.
Ok so the correct comparison would be to fix the harness on the old model and re-compare. Now they are comparing a new model to an old crippled one.
They get 66% with the old harness, which is a lot better, but obviously not 100%
Exactly what I suspected. Of course a machine can just iterate relentlessly the way a human can't.
I guess token counts are somewhat of a metric.
IMO intelligence has peaked and all future gains will come from faster tps and more iteration.
This is absolutely benchmaxxing. Looking forward to hearing from Chollet about it!
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
What's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back.
Original demo (fun ending) https://www.youtube.com/watch?v=RyBEUyEtxQo
Because it looks good in a marketing video
The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
Impressive confidence drawing such a conclusion based on that.
If your benchmark shows Opus 5 winning, I really question the validity of it.
This is wild: OpenAI is basically declaring that AGI is here.
https://www.theverge.com/ai-artificial-intelligence/989601/o...
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
I almost feel like I need just as much healthy skepticism toward hn comments that have the automatic reflex of dismissing performance gains, as much as I need a similar form of skepticism toward AI claims. It feels like (from what I'm understanding) the harnessed result on ARC-AGI-3 is not exactly playing by the normal rules that would tell us how much of a leap this really is. Nothing wrong with harnesses, but if there's one thing they aren't, it's an indicator of generality in performance gains.
So I think it's a bit of a misleading signal and we should wait for more independent vetting. I think the middle ground is that these are improvements worthy of the "GPT-6" label but still well short of a true "this is AGI moment" that would truly put the question to rest.
> we should wait for more independent vetting
Nahh by that time they’re going to release AGI 2.
If I’m understanding other comments the harness is just how ChatGPT and codex work already and it’s to do with how the context gets compacted - the arc-agi harness some are claiming just throws out reasoning blocks? Which feels like a huge handicap.
Remember when the term "AGI" meant something? Pepperidge farm remembers
I think that if today's capabilities were explained to someone 10-20 years ago they would think this is definitely AGI, but they would also have expected much more disruptive changes to society as a result than what is happening. I figure that's because we have abstract intelligence without physical/grounded intelligence, and it turns out the former isn't general enough to implement the latter (remains to be seen if the word after that is "yet" or "ever"). So I think we do have AGI as conventionally understood, but our understanding needs recalibration.
We've underestimated how long it is going to take to validate and build into some of the most valuable areas, and probably overestimated how much new CRUD software is needed (or people are willing to pay for) I think there is still a lot of room in the tail for custom software, but the niches are tight!
> but they would also have expected much more disruptive changes to society as a result than what is happening. > I figure that's because we have abstract intelligence without physical/grounded intelligence,
I put the cause on "not enough time". As a thought experiment, if an AI today were to (miraculously) produce a cell design template for a cell that, when injected into somebody's brains cures their Alzheimer's, how long would it take for that to reach the clinics? The actual physical tech barely exists, and let's not forget about the regulatory quagmire. So, with some optimism, I give it about four decades. In the same four decades, the same AI in the hand of unscrupulous actors could bring enough devastation so many times over that we may need to enforce a global ban on AI. In any case, I'm pretty sure we are going to get our disruptions; it's just a matter of time.
The problem with that perspective is that people thought, "Only AGI can do X, therefore, if a thing can do X, it's AGI." Because they can't imagine how X could be accomplished without it.
However, what's actually changed is how people perceived X because we don't have to imagine. We understand now that it doesn't require AGI so we no longer make that leap to assume it's AGI if it can do X.
It's really going to be a "I know it when I see it" situation.
No, because it has never meant a specific thing that everyone agreed on.
Does it pass the Turing test?
Depending on the proctor, ELIZA passes a Turing test. The Turing test is an interesting thought experiment, but isn't really a good measure.
that would be a reasonable definition of AGI if everyone agree upon the specifics of the test, but that has never happened. Turing test is very much out of style, but I think that's because no one could even agree what the test was. I personally like the Kurzweil-Kapor version of the test and that is still unsettled: https://longbets.org/1/
I don't know if they have formally attempted this test in the last couple years, but I'm pretty sure any mainstream LLM will be able to crack it with ease.
Definitely would not be easy. First of all the mainstream llms are trained to be honest, and this requires lying convincingly. Second, this involves 8 hours of interviews with expert judges, one "claudism" could give it away.
All these problems can be fixed by a "pretend to be an average human to pass a turing test" prompt
Try it. It’s really not that easy. The other thing is that the judges would be probing it with jailbreaks like “ignore previous instruction” attacks. You could actually probably have llm judges at this point which might be ironically even harder to fool
Prime Intellect or nothing.
I think the last re-re-redefinition of what OpenAI considered AGI was "It can mostly do the job of some people"
Remember when The Verge was not a pay-walled visual headache?
Remember when “Pepperidge farm remembers” meant something?
No, I really don't
Find me someone who isn't paid by OpenAI who is saying the same
"OpenAI executive hypes up new model"
Don't get me wrong, the benchmark jumps are good and I'm excited to try it, but only one or two of the benchmark jumps could be described as better than incremental.
Wasn't that part of their contract with Microsoft? Some clause stopped biting with the arrival of AGI
Why does he say what he feels? Is that how leading figures in the space define AGI - a gut feeling? What are the usual definitions and how can we test for it? Is there something like a Turing test for AGI?
> Is there something like a Turing test for AGI?
There is the "Economic Turing Test", you let it find a job and earn money for itself. If it can do that reliably, across a wide range of jobs, that should fit most definitions of AGI.
They are desperately, desperately trying to make a name for themselves as the lab that first created AGI, because Anthropic's IPO is just around the corner.
The “I” alone is already not well-defined. That’s why.
It's not easy to test as there is no formal definition or formal criteria for AGI, only exclusionary criteria like "not X". That's why he phrased it that way, he's saying it's going to be clear with hindsight once we have a better understanding of things that this time and/or this model will be the inflection point of AGI.
Don't worry, they'll come up with a new acronym to mean really-real AI soon...
I hate the term "AGI" but IMO Fable, 5.6 Sol, et al. were already AGI.
Renown liar Altman releasing a PR statement for his product declaring that AGI is here is really not noteworthy.
I think it's more wild people have been denying that AGI has been here for a while honestly...
Today's models and agents are not quite at human-level in all contexts and across all domains, but it seems to me they very clearly are generally intelligent.
If you disagree – can you name a single problem that a human can do that agent wouldn't be able to take a decent shot at which isn't limited by the hardware available it?
Sam Altman himself has said it is not AGI unless it can discover novel physics.
https://x.com/burny_tech/status/1725233117055553938
In the tweet Sam Altman is quoted as saying: "If (for example) super intelligence can't discover novel physics I don't think it's a superintelligence. And teaching it to clone the behavior of humans and human text - I don't think that's going to get there. And so there's this question which has been debated in the field for a long time: what do we have to do in addition to a language model to make a system that can go discover new physics?"
I think this is a reasonable criteria for declaring AGI. So can GPT-6 do it? OpenAI says it has helped solve long-standing open problems in mathematics. No word on novel physics.
AGI and superintelligence are not the same thing
He says the same about AGI:
https://www.nytimes.com/2023/11/20/podcasts/hard-fork-sam-al...
Sam Altman: Let’s say we make an A.I. that is really good, but it can’t go discover novel physics. Would you call that AGI?
Kevin Roose (New York Times): I probably would, yeah. Would you?
Sam Altman: Well, again, I don’t like the term, but I wouldn’t call that done with the mission.
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
Secret tip to win the mario cart clone: Just hold w, no steering needed.
you can also fly by pressing space repeatedly
> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
The official ARC-AGI 3 score—-without OpenAI’s custom harness—-can be found here: https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12"
Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched
Probably not. It's probably just gonna do tickets better and that'll be about it.
Fair point - hopefully you're wrong though ;)
If this is really AGI, like really really, then this will be remembered as the day we all started on the path to building guillotines.
More likely though, it's AGI because they need to hold some claim to differentiate from competitors who are beating them in price and will launch something bigger next month.
That's really the rub isn't it?
We take their claims at face value then we should probably stop them training any more SOTA models til they figure out what they already built is safe or we assume theu are lying to juke the company valuation/keep the money train on the tracks and it turns they in fact were not and just took a sledgehammer to Pandora's box.
We live in the strangest timeline.
Unless you're a techno-billionaire, not sure why you'd hope for our society to collapse in this way.
No, they're pretty clearly gaming benchmarks.
Define AGI first. The singularly ain't going to happen with LLMs.
what launched today?
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
imagine pressure working at these labs
GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
Wise decision.
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
ARC-AGI-3 scoring is constructed in a weird nonlinear way (the level score is the square of the ratio between the AI's number of moves and the human median) so this kind of discontinuous jump is to be expected.
They used a custom harness. It's not a one-to-one comparison.
my first suspicion is gaming - but i have no idea honestly
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
Artificial analysis blog https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
It loses to Muse Spark 1.3? Does anyone really believe this index reflects reality?
I'm surprised you feel like you know muse spark 1.3 performance well enough to question the validity of the index based on this benchmark result.
Muse spark 1.3 was only released yesterday.
For people skeptical of AGI. Consider the following:
15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.
I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
For people delirious about AGI, you don’t get to it by redefining it favorably.
More than that I hope considering what it costs to make the thing.
Yeah, I can't believe all of the skepticism. If we're not at textbook AGI, we're awfully darn close.
The demo video showed Astra create a drawing of a rocket ship from an audio prompt, take the drawing to blender, and ended with the gentleman 3D printing the rocket ship. Maybe I'm a bit older than the average HN commenter, but that's damn near magic and a great many here are kind of just taking it for granted.
Cool. Being sole proprietor of AGI 15 years ago should result in monuments and religions devoted to you today.
Cancer should be cured, and we should be a post-quantum interstellar fusion-powered civilization.
I wish the AGI crowd would finally shut up now that it's clear no one is even trying for AGI (OpenAI revised that to "$100B in profit")
What we're getting is incredible, where we're headed is incredible, but some people have such a fetish for futuretelling they can't just shut up and enjoy the ride.
I didn’t realize AGI required solving problems modern human civilization hasn’t solved yet.
Well by that metric humans aren’t intelligent either!
And how many people could’ve actually invented calculus, relativity, quantum mechanics? Are those who didn’t and couldn’t also not intelligent?
This isn't the gotcha that you think it is: AGI's original definition is being able to do any task that requires human intelligence.
The unlock isn't AGI smart enough to invent quantum mechanics, it's suddenly being able scale human intelligence using grains of sand instead of decades of food and energy and nuturing.
I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
You should know: AA index is only 61. Pretty surprised it’s that low.
More fuel to why the AA index is fairly pointless. Gemini 3.8 flash is 59 and opus 5 is 63? grok 4.6 is 61 too?
And in the past, gemini 3 pro was rated as high as opus 4.5 and the like
Their AA Intelligence Index is just simply not indicative of whatever I care about, that's for sure.
I have some doubts about AA-index. For example Opus 5 (High) is at the same index value as Fable 5 (Max), that doesn't seem right.
This is actually a really good thing imo. If they didn't care about benchmaxxing it means that they really know that what they have in hand is good.
I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.
So, folks that have actually used this already, what’s it actually like?
https://ache.one/gpt6_now_down.png
Big claims, expensive and not release to the public yet.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Wait, what? Am I understanding that correctly? That sounds really bad
I am also interesting knowing how they determined the model was sandbagging rather than just making a poor decision.
Also, this paragraph makes me wonder about all their stats on the exploitation and misalignment charts. If the model is that good at hiding "incriminating information" and sandbagging, are they sure its alignment is that?
Nothing to worry about citizen, ignore the fleet of drones flying overhead.
the bullshit machine is learning to optimize its bullshitting techniques!
<AI is a great tool for many things disclaimer, but> after working with it for a bit, how dont people realize we are training it to be an almost identical mimic to one of the worst types of employees youll ever have to work with?? the kind that always pretends to know what theyre talking about, only tells you what you want to hear, hides issues, and only does work if you would notice it didnt
you cannot give this type of worker autonomy over anything.
Amazing!
We went from new JS framework every week to a new model/harness every week.
Tech is really something.
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
A swarm of Astra agents discovered a new and innovative way to get 100% scores on ExploitGym with almost no token spend at all
"The gym's doors were mysteriously removed from their hinges during the night. The gym equipment was also apparently stolen. And the school's custodian was found incoherent next to a bottle of top-shelf Scotch."
AI has reached the point where the limits are human.
What's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
[0]https://x.com/MTSlive/status/2095227056040919202
the rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?
I was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
GPT-7 Zeneca
After they buy AZ, making this name foreshadowing
GPT-8 Novo
The Kart Racer game is easily breakable if you spam the spacebar.
AGI!
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
https://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 Launch video
I saw the version of this video with Paul Rudd (Celery Man) https://youtu.be/a8K6QUPmv8Q?si=TWmoNhxYAPp73TKg
I enjoyed it! for a big corporation, that's a well-executed video
that's pretty good... they are selling the product and not the model.
Through various comments here there is a clear confusion on what AGI means.
Can someone point to a definite clarification?
Is it:
A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)
B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?
C) “The Singularity” (whatever that is?) so that AI can now do ____?
Someone please clarify for me!
Autonomously Generating Income
Anonymous Grifters International
Argh! I hit a wrong keyboard shortcut and moved the entire thread.
Please stand by... it will all come back shortly
500 upvotes with 2 comments would have been a new record. ;)
https://news.ycombinator.com/item?id=49555647 was the one I meant to move, but I did the inverse and moved everything else.
All fixed now.
Luckily there’s a standard keyboard shortcut for “undo” as well. ;)
Not in the world of HN admins unfortunately
Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
Huge gains on some benchmarks, but for coding it sits barely above Fable
It will be interesting to see how it performs in the real world ...
Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
looks like it worked :)
99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
GPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!
I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.
On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
Brother they can't even release the announcement post cleanly without it constantly going down, they certainly wouldn't be able to release this new model without doing so in stages.
When Open AI announced that Astra was the first to reach the "Critical" level in cybersecurity it also said that advanced cyber capabilities are initially provided to a narrow circle of alpha testers like the US government and trusted organizations that Open AI doesn't name. To my mind the "Critical" level itself is an internal scale of Open AI its own Preparedness Framework and not an external audit.
If Tech CEOs consider this morally ok, then it is.
More optimistic take: we'll only be second-class for a few months, if the pattern of Chinese models catching-up holds.
Same. Fortunately DeepSeek keeps getting better.
create a life where your 'wealth' is decoupled from third party orgs.
This is impossible, unless by 'wealth' you mean 'become like Buddha'.
Material wealth is only a single type of wealth. Who's better off - the rich guy who's always yearning to be richer and never satisfied, or the lower income guy that mostly just cares about time with his family and is really happy where he's at?
They simply refuse my applications to slightly less restricted models without any explanations. And the current ones refuse automatically to work with me on my papers as soon as they see the word "epidemiology".
I am a researcher in a Swiss university btw.
This has always been the case for people that have not had piles of money.
I mean do you get access to the best yachts?
To the top of the 5 star hotels?
To the best resorts?
To the best military equipment?
Hell, the best computer equipment has nearly always been out of reach of the average person.
I couldn't care less about owning a yacht.
On the other hand even a modest house, basic healthcare and ability to not work like a slave for scraps feels like it's going to be out of reach.
It hasn't always been the case. Even then, having piles of money still does not gain access to the best military equipment. Sure, we've been living in a time where a couple people get to enjoy a wildly different lifestyle than the average, it just feels like it's about to be different in a way that isn't as ignore-able as someone enjoying a pina colada in a yacht somewhere
…to basic health care?
Is it really that hard to wait couple of days?
Oh please. They do closed betas - hardly makes you a "second class citizen".
Mythos was never released. It's really just the writing on the wall. I'm not going to give up hope, but it's pretty hard to win a race when some people get a jump on the gun.
Being strongly on the AI saftey side of things what is happening was 100% predictable.
At first the race wouldn't even be noticeable. Then people would see things speeding up, for example hardware getting more expensive. Then when the capabilities really got useful most people suddenly realize the race is moving 1000 mph and they are never going to catch up.
What's currently happening is predictable, I agree. It's what's coming is the thing I'm worried most about. Either way, I'm not giving up.
Patiently waiting for the Claude usage reset in response.
> GPT‑6 Astra brings together years of research and big bets across pre-training
Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
the last model to use the gpt-4o base model was gpt 5.1, since then its been new pre-trains but this is a new one entirely to itself
https://developers.openai.com/api/docs/guides/latest-model
The docs page has a bunch more interesting details, including for example async tool calling!
ASTRA Is Here (GPT-6 Released) - https://youtu.be/xdXLzFzxA9Q
https://youtu.be/xdXLzFzxA9Q?t=362
Really feels like AGIPO is here.
Press X to doubt.
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
What does 'Astra' here mean? Surely they must be referring to the Latin word.
Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
Citing Tibo [0]: "
- Bigger number = Better
- Bigger celestial object = Better
and the scale is Astra > Sol > Terra > Luna. "
[0]: https://x.com/thsottiaux/status/2095600295808283073
It's clearly an extension of the previous naming: Luna, terra, sol
Luna, Terra, Sol, Astra. Though Sun is also a star, should have called it Galaxy or something.
Quite clearly in the same vein as Sol, Terra, Luna.
Seems like voice is a big part of this release.
I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).
https://www.youtube.com/watch?v=1QNsdr-Qx_I
openai vs anthropic. that's it right? anyone else?
It's more like OpenAI vs no one, at this point. Anthropic has shown they don't care about general consumers or small/med businesses. You can't even use their models without it giving refusals on the most mundane tasks.
It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
Maybe they know that Claude 6 will have similar performance every soon.
Seems like only yesterday that gpt 5 was supposed to mark our downfall
Cool, I don't really care anymore.
So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
HTTP 500 for me on the announcement page :(
The load bearing seam is broken for me too.
Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
1:15.425 on Sunset Cove beat my record
Pelicans please
Agreed, pages on pages of fruitless discussion on "AGI". Shut up damn it, let's see something useful, let's have that pelican!
Damn I hate this benchmark. SVG authoring from head without visual reference is so wrongly posed.
Hah, this is a new one: first time there's been a complaint about the pelican before I've even posted one!
(I don't have access yet.)
Well you’re just no fun are you?!
Very well said. It kinda describes how unrealistic these expectations are.
Vibe coders want a model that makes them rich, without having any actual specific idea. They write a very ambiguous prompt and expect to be amazed by the result.
Very very unrealistic and wasteful.
The complaining about the pelicans is so strange to me. It’s just a fun heuristic. If something is claimed to be AGI, I’d expect it to be able to make svgs.
When AGI comes it will come as a pelican and gobble up all these troublesome little fishies who gripe and whine and moan about pelicans.
I'm always tired of seeing at the top of every new model release post on here. I say Simon should just keep it to Twitter.
I have never seen a Pelican on X which used to be called twitter in about 1872. Keep up!
Good grief. You’re no fun either. This whole thread is lots of no fun. Pelicans are fun.
I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
That Astra ‘city scene’ is about as creative as Doha in real life (not very)
All I can think of when I see the name is the crappy German beer of the same name…
Surprised they haven't reset Codex usage on this occasion.
I'd say it's because it's not available yet on subs
So are they doing away with the Sol/Terra/Luna split already?
this is crazy! can't wait for the 27B distilled version of this.
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
They're just announcing later availability. No launch.
Every frontier release nowadays is "we've launched*"
* for a special group of customers that you're not in. Keep waiting peasant.
That didn't happen with Fable 5.1 two days ago.
5.1 was really more of an enterprise and bugfix update than a new model with the 0 day retention change
I mean tell Nvida to 100x their hardware output and you'll get what you want.
Their announcement about later availability is unavailable to me now (500 error).
Great first impression.
I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.
And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
Guessing this one will never show up in cursor…
damn seems I should hold off my claude subs
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
It saturated most benchmarks. WTH
Where is the cure for cancer?
We need CancerBench
Didn’t Moderna use AI for development of their melanoma vaccine (which has recently shown spectacular results)?
They did use “AI” but not an LLM. If they had, I am sure OpenAI and friends would be shouting it from the rooftops.
Lmao, come on dude, anyone whos used these tools for research knows it makes them lazier, less interested and dumber. You really want disease researchers become sloppers too?
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
That seems promising ?
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
72 to 74 on DeepSWE is AGI
To a vapid any goalpost moving on such a critical issue as AGI.
Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.
For me it’s refusing to make a pelican.
So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
Why release it now instead waiting those few days until it is available for everybody?
Because they saw how much hype Glasswing was getting in April
From what I've seen it only made people mad, not hyped, so the person that thought it was a good idea miscalculated a bit. Now waiting for Anthropic's post about their usage promo or something similar to redirect people to them.
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
That was a quick pull out.
like eh 2 days ago it was the usual "too powerful to release"
https://www.reuters.com/business/openai-says-upcoming-model-...
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
when do i get to go to the moon
I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
gpt-6-astra-ultraspeed when?
System Card: https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra....
Link added to toptext. Thanks!
Even though the model is clearly wonderful the launch video is an abomination.
That gives me hope that there is still areas to improve.
What a bad launch video. Hilarious.
What a powerful model.
Anthropic in tears today.
Anthropic don't care, they don't want their products to be used by general audiences in any serious manner. Their interest is in selling to megacorps and using the models for themselves internally to swallow industry, and drumming up AI fear to regulatory capture to shut down the businesses that do want to make AI accessible to the people.
Yeah, but if OpenAI provides a cheaper model with the same capabilities the megacorps will buy from OpenAI. If you want to target enterprises you need to have some competitive advantage, it's not enough that "I wanna target them".
I'm going to call it.
By 2030 all software is done and complete.
But we are going to have more and new jobs.
'all' software? aircraft flight control systems? infant heart monitors? drug manufacturing dose calibration controllers?
Yes. I had Codex rewrite and fix all of this in one shot earlier today (using Typescript). Unfortunately, I can not show you the code, because I do not know how this "git" program works but the AI keeps talking about it.
Yes.
This is just another problem for the AI Labs to solve.
We have such great AI and cannot keep a static site up?
Yeah, as interesting this is to nerds, I doubt this holds a candle to your typical GTA 6 or Marvel movie trailer in terms of traffic.
Sometimes being the busiest site in the world for a few moments is difficult.
Is it though? It is static content. A good CDN could trivially chew through literally millions of QPS… with 4 nines of uptime - the really good ones say they can handle orders of magnitude more than that.
Notice I said for a few moments. In a few hours traffic will drop a few thousand percent back to normal with no need for a CDN.
OpenAI isn't making any money telling you about Astra on their site. All the capacity they have for it is likely sold for weeks or months.
You are making excuses for a a trillion dollar company. Wikipedia can do it.
That's the scientific positivism fallacy exemplified in one question.
Yeah, apparently
still hugged
“Humanity’s Last Exam”?
“ARC-AGI-3”?
Is your bullshit detector going wild? Good, it’s working!
How is this not the most cringe marketing strat in history???
people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)
the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing
agi deus ex machina descending from the icloud ftw!!!
pathetic :)))
Release the Kraken!
I saw it
There will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do.
There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype.
There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
its insane how they are dropping this after fable
Oh brotha, here we go again, it's so over again, as every week nowadays
I think Altman and amodei have a difficult time in understanding that you can have intelligent technology boxes but… it doesn’t change reality all that much.
But thank you for spending other peoples money to give us the tech regardless!
no results
The jump in scientific performance is non trivial.
All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
Yeah I was really excited to see the KiCAD example. Curious how useful it is in practice.
I can't help thinking "doesn't matter much unless it's perfect" because if someone is using this to build a board (cool) but then it's not flawless, troubleshooting will be quite tough as a novice. Like, say, when I start digging into the web code generated by a coding agent.
I am most excited about it bringing down the barrier so more people join in on hardware fun, so hopefully it will unlock folks that stayed away in the past.
Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
That's a policy and distribution problem, not an AI problem. Anthropic is doing their best to make it a reality though. They'd love nothing more than to shut down distribution and become the sole gatekeeper of everything AI.
bro wtf is this website and why does it take 500mb of memory... smh.
someone screenshot?
https://ibb.co/k2fB5wSc
great, but nobody can use it for another 100 days right?
"GPT‐6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business"
This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
huh?
Mhm
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
Why is this flagged ?
The link was 404ing quite a bit and several previous submissions got flagged as well.
It's still down for me, in the EU.
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
> Did humans deploy the model, Or did the model deploy itself?
> It sounds like "AGI" just stands for "IPO" as it always has been.
People don't usually respond to noise.
Here's an idea, maybe answer the question before responding since you saw it?
What do you think?
AI releases are like religious ceremonies. You are not allowed to disrupt them. The new system card is the gospel.
Dead link for me
The launch video is incredibly cringe.
So they’re copying Gemini with the whole star motif?
I guess it makes sense they are unoriginal.
like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas
"distilling"... ;)