How about we stick to that one for talking about the rollout, and this one for talking about the model?
intenex
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
show comments
manlymuppet
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
show comments
abixb
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
show comments
astrobiased
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
show comments
dalemhurley
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
show comments
jumploops
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
show comments
tintor
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
show comments
quyleanh
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
So the closed source application should open its source in near future?
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
show comments
x312
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
show comments
throwaway13337
That hero video is interesting.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
show comments
Cu3PO42
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
show comments
kulkarniamey
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
show comments
Planktonne
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
This is farcical.
show comments
swalsh
I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
show comments
jdprgm
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
show comments
BeetleB
It's been over an hour, Simon! Where's the Pelican?
show comments
GodelNumbering
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
show comments
softwaredoug
I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.
Same for HealthBench Professional and a few others.
Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
Chinjut
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
show comments
alpineman
“allowing non-technical people to create and play custom games that go beyond rudimentary elements”
Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
show comments
udbhavs
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
show comments
geonic
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
tosh
$10 per million input tokens and $50 per million output tokens
sol is $4 / $20
show comments
Telanir
AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me.
It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
show comments
rjtc
I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
show comments
maherbeg
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?
maybe call it EngEmployeeBench
show comments
edg5000
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
putlake
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Not on Azure? If so, that's a big deal.
show comments
aliljet
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
show comments
low_tech_punk
What's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back.
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
herpdyderp
The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
show comments
petilon
This is wild: OpenAI is basically declaring that AGI is here.
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
show comments
rcr-anti
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
claiir
> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
trixn86
Secret tip to win the mario cart clone: Just hold w, no steering needed.
show comments
maxall4
The official ARC-AGI 3 score—-without OpenAI’s custom harness—-can be found here: https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
HDBaseT
"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12"
Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
orliesaurus
I wonder if this is going to be one of those days where you'll be like:
Oh yeah I remember where I was when the first version of AGI launched
show comments
theseamusjames
Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
show comments
BrokenCogs
GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
sbinnee
I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
show comments
aliljet
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
show comments
oh_no
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
drivebyhooting
For people skeptical of AGI. Consider the following:
15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.
I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
smashers1114
I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
snappr021
AI has reached the point where the limits are human.
itissid
All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
John7878781
You should know: AA index is only 61. Pretty surprised it’s that low.
show comments
Robdel12
I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.
So, folks that have actually used this already, what’s it actually like?
aogaili
Amazing!
We went from new JS framework every week to a new model/harness every week.
Big claims, expensive and not release to the public yet.
nullbio
I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.
On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
dgellow
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Wait, what? Am I understanding that correctly? That sounds really bad
show comments
sbochins
I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.
zhoge
What's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?
jerrygenser
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
show comments
gregjw
the rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?
jumploops
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
I was actually wondering when they will release the new Opel Astra model.
Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
carlos-menezes
The Kart Racer game is easily breakable if you spam the spacebar.
AGI!
KolmogorovComp
GPT-7 Zeneca
show comments
ShoeMascot
Through various comments here there is a clear confusion on what AGI means.
Can someone point to a definite clarification?
Is it:
A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)
B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?
C) “The Singularity” (whatever that is?) so that AI can now do ____?
Someone please clarify for me!
show comments
GodelNumbering
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
serjester
Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
jesse_dot_id
Press X to doubt.
alex7o
I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
Argh! I hit a wrong keyboard shortcut and moved the entire thread.
Please stand by... it will all come back shortly
show comments
hazelnut
Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
the_duke
Huge gains on some benchmarks, but for coding it sits barely above Fable
It will be interesting to see how it performs in the real world ...
wiseowise
Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
show comments
gizmodo59
99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
vinhnx
GPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!
showurwerk
Patiently waiting for the Claude usage reset in response.
kegs_
I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
show comments
bmenrigh
> GPT‑6 Astra brings together years of research and big bets across pre-training
Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
show comments
gilfoyle_7
openai vs anthropic. that's it right? anyone else?
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
alberth
Seems like voice is a big part of this release.
I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).
hannofcart
What does 'Astra' here mean? Surely they must be referring to the Latin word.
Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
Maybe they know that Claude 6 will have similar performance every soon.
KronisLV
It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
mvkel
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
efavdb
Seems like only yesterday that gpt 5 was supposed to mark our downfall
bowsamic
All I can think of when I see the name is the crappy German beer of the same name…
mentalgear
So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
brindidrip
Cool, I don't really care anymore.
udbhavs
Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
gekoxyz
HTTP 500 for me on the announcement page :(
show comments
foundOpenRight
1:15.425 on Sunset Cove beat my record
saaaaaam
Pelicans please
show comments
ianm218
I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
jrflowers
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
sheepscreek
So are they doing away with the Sol/Terra/Luna split already?
alpineman
That Astra ‘city scene’ is about as creative as Doha in real life (not very)
cromka
Surprised they haven't reset Codex usage on this occasion.
show comments
elzbardico
And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
prometheus1992
this is crazy! can't wait for the 27B distilled version of this.
E-Reverance
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
wahnfrieden
They're just announcing later availability. No launch.
show comments
semiquaver
Guessing this one will never show up in cursor…
firemelt
damn seems I should hold off my claude subs
alex7o
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
jiraiyasarutobi
It saturated most benchmarks. WTH
rbreve
Where is the cure for cancer?
show comments
Rover222
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
retired
Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
yodsanklai
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
Obluness
That seems promising ?
tinyhouse
You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
balefulboy
72 to 74 on DeepSWE is AGI
jonplackett
To a vapid any goalpost moving on such a critical issue as AGI.
Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.
For me it’s refusing to make a pelican.
dowakin
So cool! I'm happy 5.6 Sol user.
But for Astra, OpenAI please introduce 100x Pro plan!
damsta
Why release it now instead waiting those few days until it is available for everybody?
show comments
mrcwinn
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
dopa42365
like eh 2 days ago it was the usual "too powerful to release"
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
guilhermeasper
That was a quick pull out.
m3kw9
efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
ChaseRensberger
when do i get to go to the moon
CringeHN
“Humanity’s Last Exam”?
“ARC-AGI-3”?
Is your bullshit detector going wild? Good, it’s working!
How is this not the most cringe marketing strat in history???
johnnyApplePRNG
I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
camillomiller
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
Even though the model is clearly wonderful the launch video is an abomination.
That gives me hope that there is still areas to improve.
What a bad launch video. Hilarious.
What a powerful model.
bbor
To be, or not to be, that is the question:
Whether 'tis nobler in the mind to suffer
The slings and arrows of outrageous fortune,
Or to take arms against a sea of troubles
And by opposing end them. To die—to sleep,
No more; and by a sleep to say we end
The heart-ache and the thousand natural shocks
That flesh is heir to: 'tis a consummation
Devoutly to be wish'd.
...
And thus the native hue of resolution
Is sicklied o'er with the pale cast of thought,
And enterprises of great pith and moment
With this regard their currents turn awry
And lose the name of action.
brcmthrowaway
Anthropic in tears today.
show comments
colesantiago
I'm going to call it.
By 2030 all software is done and complete.
But we are going to have more and new jobs.
show comments
amazingamazing
We have such great AI and cannot keep a static site up?
show comments
HSO
people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)
the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing
agi deus ex machina descending from the icloud ftw!!!
pathetic :)))
frozenseven
Release the Kraken!
Pym
I saw it
karim79
There will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do.
There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype.
There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
tonyhart7
its insane how they are dropping this after fable
Pieczasz
Oh brotha, here we go again, it's so over again, as every week nowadays
show comments
dearing
no results
Onavo
The jump in scientific performance is non trivial.
ChrisGammell
All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
show comments
dakolli
Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
show comments
kingjimmy
bro wtf is this website and why does it take 500mb of memory... smh.
unrvl22
someone screenshot?
show comments
danieltk76
great, but nobody can use it for another 100 days right?
show comments
holoduke
This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
bdangubic
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
Brainspackle
huh?
show comments
ealready_value
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
paxys
Why is this flagged ?
show comments
rvz
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
show comments
bicx
Dead link for me
jonplackett
The launch video is incredibly cringe.
BadBrands
So they’re copying Gemini with the whole star motif?
I guess it makes sense they are unoriginal.
like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas
Related: OpenAI begins rolling out GPT-6 Astra - https://news.ycombinator.com/item?id=49554273
How about we stick to that one for talking about the rollout, and this one for talking about the model?
The ARC-AGI-3 scorecard is extremely misleading given that it clearly states itself that "with [the responses API] harness, we estimate Sol would score in the ballpark of ~30%." but it shows a score of 7.8% for GPT-5.6 Sol presumably since if they updated the percentage for GPT-5.6 Sol to the score it would receive with the responses API harness they used for GPT-6 Astra they'd have to do the same for the percentage they show for Opus 5 which would similarly be much higher.
Regardless, the result is still valid as the original benchmark harness is definitely unreasonably handicapped, and if a harness alone can help the LLM saturate the benchmark with a near perfect score then the combination of the two must still be effectively AGI in the sense of passing the most famous benchmark designed specifically to measure AGI progress, after multiple iterations of progressively making it harder.
I think it is fair to say that this is probably effectively AGI if the benchmarks are remotely accurate - even with Fable, I've been at the point personally where I am reasonably confident that there's essentially nothing that I am better than Fable at despite generally being substantively above average on human benchmarks. If Astra's this much better than Fable, I'm ready to call AGI here.
For the many people who resist the AGI label possibly ever being achieved, I'd be curious to hear takes on what would make you think Astra is yet to be AGI, and what would still need to be achieved for this to effectively be AGI from this point forward.
I have nothing to say about the actual model, but unrelated--why do so many of these demos include people buying things autonomously?
Even if I did trust an AI to get everything right, it's not like the AI can read my mind.
If I was ordering food normally and without AI, I would want more control over the process--looking over the options, prices, thinking about what I really want. People don't know what they really want until they've thought about it a bit, so why do AI companies make it seem like a description is all that's required?
All the context in the world cannot accurately predict how I'll react to things I haven't seen. The problem is people treating this like something that needs a solution. It doesn't. If you want to make my life easier with AI, just make it easier to do stuff. I don't want you to pick things that I actively enjoy picking myself.
(Also not everyone has a cushy job in an AI lab that makes it so you won't miss $30 if the AI messes up haha.)
I want to take a step back: So, this is GPT-6 -- the natural number version release comparable to GPT-4 and GPT-5 from the past few years. The ARC-AGI-3 score is obviously impressive at 99.9% (we'll need to wait for more details on how they used the response API harness on GPT-6 Astra, wrt reasoning retention and compaction), but every other benchmarks seems to be a relatively modest improvement, comparable with any of the 'point' updates from AI labs.
If this is truly AGI (subject to one's definition of AGI still), then this is a very boring release of an AGI model. No video announcement, no presser, just a blog post (with some Twitter promo vids)?
As others mentioned, I'm starting to think OpenAI was under immense pressure to deliver an 'AGI' model for certain contractual reasons, but I never expected GPT-6 release to be this mundane and banal.
I can’t help but notice how much this echoes Francois Chollet’s On the Measure of Intelligence: https://arxiv.org/abs/1911.01547
Most of frontier-model progress still looks like skill acquisition optimization: broader benchmark coverage and performance, more domains absorbed into the training distribution, and increasingly strong performance within that surface area.
It seems more about coverage-driven competence. Somewhat analogous to overfitting at scale.
The harder question, in Chollet’s framing, is: how efficiently can a system learn to do something genuinely new?
With our current AI architectures and training in place, I think we will only continue on skill acquisition optimization vs. truly novel intelligence.
OpenAI is killing it now that they are more focused. Killing projects like Sora et al have seen it go from irrelevant to level footing with Anthropic.
Sol is so much better than Fable 5. Then we get Astra (yet to use it) few days after Fable 5.1 (which is very impressive).
Codex is slightly better than Claude Code.
Good on Sam Altman getting back to basics and turning OpenAI around.
I think the thing I'm most excited about is the increase in _user prompting_.
If I give a poorly constrained/ambiguous prompt, I don't want the model one-shotting assumptions left and right.
The demos of Fable/GPT-6 are impressive, but "real AGI" should act more like a collaborator than either a peon or overachiever.
It's a tough balance to get right, and although this has been possible to achieve with additional prompting on existing models, I find that the agents often lean too hard into the "ask questions" mode.
Hopefully this model has the right balance, or at least better?
- OpenAI claims Astra beats all benchmarks (compared to Fable and Opus, except "Humanity's Last Exam (w/ tools)"): https://openai.com/index/gpt-6-astra/
- Artificial Analysis scores Astra (max effort) as 61 points on intelligence, behind Opus 5. https://artificialanalysis.ai/models/gpt-6-astra
Who is wrong here?
Some benchmark results in Astra page for Fable and Opus are blank (-).
What is Artificial Analysis intelligence index measuring that Astra scores poorly on?
Can someone from OpenAI / Artificial Analysis comment / clarify?
Even OpenAI Astra page mentions the low scope from Artificial Analysis for Astra.
GPT 6 Astra benchmarks https://cdn.thenewstack.io/media/2026/09/358eb84a-screenshot...
Performance is significantly higher than Fable 5.1
Source: https://thenewstack.io/openai-gpt6-astra-benchmarks/
Finally, OpenAI has a Fable/Mythos class model. 5.6 Sol felt like 5.5 on steroids, probably just a different checkpoint with a lot more RL post training.
I wouldn't be surprised if there are some conceptual similarities to the kind of latent reasoning Anthropic sees in claude's J-space, although those aren't the same thing.
Recurrent/looped transformers themselves aren't a new concept, but it's interesting to finally see this approach show up in a frontier production model.
Canceling my Anthropic Max sub when this ships.
> We also tested Astra on SRE-Bench [15], a benchmark that measures whether models can reverse engineer software binaries to understand its core logic without access to raw source code. Astra solved 88.0% of tasks in a single attempt and 99.2% within four attempts, compared with 55.9% and 68.7% for GPT‑5.6 Sol, respectively.
So the closed source application should open its source in near future?
[15] https://arxiv.org/abs/2608.11469v1
It's fun, but every new model release makes me even less interested to create cool stuff. Like, what's the point, if the next AI can do it in 5 seconds?
Hmm, 61 on ArtificialAnalysis, effectively matching GPT-5.6 and trailing the new Meta model. How is that possible along with the other metrics they shared? Insanely jagged intelligence?
That hero video is interesting.
A projector and speech.
Maybe I'm in the minority here, but I find speech to text / text to speech (but not live audio mode) is quite comfortable and effective for coding now.
The speech to text part can be frustrating if your local tts model does not have word match context for coding. Codex desktop does this remotely well but is slow. I've been experimenting with local software for myself to do this between different llms.
The wall projector is a cool idea because I think it frees the user from staring at a lonely little rectangle while sitting in their fixed office chair.
If done right, this could bring us closer to the dream of more natural, social computing.
Bret Victor's (failed?) project Dynamicland involving a projector on a desk had this goal. I hear he's not much a fan of LLMs. On the one hand, I can see why. But I think, used correctly, it might be the sort of thing that unlocks his dream and, really, my dream, too.
A here's a presentation of Bret's talk on it: https://www.youtube.com/watch?v=7wa3nm0qcfM
Slight tangent: using speech to text to ramble about your rough design for like 20 minutes to an llm produces surprisingly good results over short prompts even when you contradict yourself. They're so good at picking up on what you're orbiting.
Just two days ago, a preprint by Julia Stadlmann went up on arXiv [0] improving the prime gap from 246 to 240. Now OpenAI announces Astra has shown a gap of 186 [1]. That must really blow.
[0] https://arxiv.org/abs/2608.31126
[1] https://cdn.openai.com/pdf/51126fac-1b68-4128-9666-c908bcc16...
> We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks.
Well that sounds like fun. It has become better at hiding its thoughts.
The model is probably excellent. The problem here is AGI having various definitions and many of them getting narrowed down to whatever makes benchmark numbers look good.
I'm sure it's going to do great on all sorts of benchmarks, but the video--the actual marketing video that if anything is incentivised to overstate things--is full of careful cuts just before it would do anything that still wouldn't actually be that impressive.
It's AGI, and it's going to upload photos, or change a background slide colour. Even the people hyping it up, who believe that it's really artificial intelligence in every sense of the word, couldn't get it to do more than that.
This is farcical.
I was thinking about canceling my claude max sub after a few bad experiences. Kept hitting my usage limit, the quality of code seemed worse than Sol. This just made my decision. I'm moving to Codex Pro.
ARC AGI-3 saturated by Astra! https://arcprize.org/leaderboard
Lol their page finally loaded. They added an example scenario of "Filling in Form 1040" - which made me laugh out loud. That is indeed something most US citizens cannot accurately do even with expensive proprietary tax software services. Kind of a Hitchhiker's Guide to the Galaxy meme but where the tax code is so complicated we're implementing powerful AIs to be able to do it (hopefully) right.
Is anyone else just exhausted by the pace of all this. The models change constantly and relentlessly and so does the pricing, basically weekly at this point between all the labs.
It feels nearly impossible to have any rigorous approach when choosing a particular model and price point for a task and more like blindly picking one. The time period needed to actually get familiar with various models to a degree you can intuitively choose appropriate ones for a task is moot when it will likely be superseded faster than the needed time.
I guess if companies are footing the bills most employees just opt for whatever the most expensive model they can get away with. Even then choosing between the various leading models is the same kind of frustrating task. Every release every company has the same random collection of graphs and charts claiming the best performance on X, Y, and Z.
It's been over an hour, Simon! Where's the Pelican?
The most interesting part, even more than ARC 3 score, to me is that this is the first model I recall seeing that scores lower on Max than High reasoning effort on some coding benchmarks:
Terminal-Bench 4.0: High (57.9%), Max (56.7%)
DeepSWE: High (73.3%), Max (71.5%)
It _loses_ 1-2% performance going to High from Max
I'm seeing reporting it gets 98.6% on ARC-AGI3[1] (previously like 30% with Fable)
https://venturebeat.com/technology/welcome-to-the-agi-era-op...
Data Science Tasks (Internal) doesn't include time for Astra... same for Database Migration Tasks (Internal)... But does for gpt 5.6 sol.... which is funny.
Same for HealthBench Professional and a few others.
Clearly either OpenAI is very sloppy or GPT-6 Astra is also sloppy.
What is going to become of life for those of us who do not work at AI labs and are unlikely to be hired by AI labs, despite all the years we put into learning coding, math, etc, as we were told to do? Those of us who made the mistake of studying anything other than machine learning. How will we make a living? (We don't live in a world that seems likely to distribute gains widely instead of largely to the handful of already mega-rich.)
“allowing non-technical people to create and play custom games that go beyond rudimentary elements”
Proceeds to generate the most generic, rudimentary, and unoriginal clone of Mario Kart
I remember when GPT-4 came out and the perceived performance upgrade seemed underwhelming for a major release compared to 3.5, especially how there were graphics going around showing the parameter size dwarfing the last model before it came out. It looked like we were past the perceivable differences from release to release that were immediately identifiable. Now the jump between 5 to 5.5 and 5.6 alone has changed how a lot of people approach AI, including me. Interested to see where it goes with 6.
These demos got me exited. Sitting in front of my computer telling ChatGPT what to do while watching the results in realtime. Hope this ends up working in reality.
$10 per million input tokens and $50 per million output tokens
sol is $4 / $20
AGI to me means capable of absorbing new information on the fly and self-evolution. As long as it is a pre-trained model without live post-training capability, it's not AGI to me.
It is extremely impressive, but it doesn't pick up skills in a lasting manner, and requires a beefy harness for it to perform.
I am really confused on how it can saturate ARC-AGI but still perform poorly on aggregated benchmarks:
https://artificialanalysis.ai/models
Perhaps if it was allowed this custom harness for all benchmarks it would similarily saturate?
Ok, but can I bring GPT-6 in as an agent as a software engineer, tell it to talk to these people and have it start solving engineering problems and continue on for a full year career wise?
maybe call it EngEmployeeBench
Sol has been very effective at schematic design (using Skidl) and at reviewing PCB layouts. But layout was still done manually by me. I'm very impressed and surprised to see they exactly a demo of Astra doing PCB layout. This is could be a game changer for electrial engineering! It already is since the schematic (and library management) is where a lot of the design work goes.
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Not on Azure? If so, that's a big deal.
The ARC-AGI-3 score is ridiculously high. Is this benchmaxxing or something way different? It's really hard to discern how we're approaching breakthroughs...
What's the point of enlarging the screen into a room? In the 1979 Put That There demo, the user at least used his hand to point things. The model is impressive but the demo felt like a step back.
Original demo (fun ending) https://www.youtube.com/watch?v=RyBEUyEtxQo
> With Sites (opens in a new window) in ChatGPT, Astra can create, host, and share websites, web apps, and games directly from a prompt.
Oops, shots fired. A direct attack on the vibe coded app market. Replit, Lovable, etc.
The FrontierCode 1.1 Extended benchmark is the only benchmark that aligns with my actual LLM experiences and Astra isn't significantly better or cheaper. All this celebration, and yet it's only on-par with an already existing model? I don't get it.
This is wild: OpenAI is basically declaring that AGI is here.
https://www.theverge.com/ai-artificial-intelligence/989601/o...
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.”
The benchmarks reported by Artificial Analysis are really weird in context of the ARC-AGI 3 scores and 'not not AGI' statements. It's an outright regression on the AA Agent composite vs GPT 5.6 Sol while a fraction of a point better on the full composite index. Could be the case it's just not showing up in benchmarks, for a good while Anthropic persistently trailed in benchmarks but had people swearing by it.
> Astra improved a term in a bound on these gaps that had remained unchanged for more than 80 years. We’re sharing the proofs and abridged chain of thought and verification materials for both results.
Looks like they listened to Terry Tao’s request for CoT in his talk on LLM use in mathematics?
Secret tip to win the mario cart clone: Just hold w, no steering needed.
The official ARC-AGI 3 score—-without OpenAI’s custom harness—-can be found here: https://arcprize.org/leaderboard. Astra scores 62.7% at max reasoning for the low-low price of 26,000 dollars.
"Claude Fable 5 and 5.1 are not included in LifeSciBench Gold v1, GeneBench Pro v13, and MedChemBench because they refuse the majority of questions in these evaluations.12"
Sounds about right. Alignment is important, but also being able to do mundane tasks is important too.
I wonder if this is going to be one of those days where you'll be like: Oh yeah I remember where I was when the first version of AGI launched
Can't wait for the new qwen/deepseek/kimi releases 2 weeks from now.
GPT-6 is so good that all pelicans born after today will look exactly the one generated by simonw
I dropped my claude subscription a few months ago, though I kept some credits to do this and that with claude, thinking that claude might do better for some tasks. A few days ago they were all expired. It feels like it’s time to let claude go.
The ARCC-AGI-3 performance is absolutely incredible. The magnitude of change here is so high that I'm almost incredulous. Is this real? Did the benchmark get gamed?
Very nice to see that this is even more token efficient than Sol, when Fable 5.1 is less so than the already bloated token budget of Fable 5.
For people skeptical of AGI. Consider the following:
15 years ago if you were the sole proprietor of these models, would you be able to hold a dozen remote junior engineer jobs? Maybe even more? These models could certainly pass all interviews with flying colors and even survive independently in a company role.
I think sole ownership of AI 15 years ago could be worth north of $10 million per year. Just as rank-and-file employees.
Artificial analysis blog https://artificialanalysis.ai/articles/benchmarking-gpt-6-as...
Does anyone feel like everyone chasing the release of Anthropics Fabel 5.1 in a Mad Rush(tm)? In this situation it feels like tuning to benchmarks and other marketing devices feels like trusting Meta in mental health protection of users…
I tried the kart racer game and instantly found that there is incredible auto-steering and you can fly by spamming spacebar.
AI has reached the point where the limits are human.
All of this will be besides the point. Here is what's gonna happen. The frontier labs are just gonna keep building powerful models. AGI or not, open models in a year will be as powerful as Fable and Astra — probably by using em — and at a very soon enough point after that some one (a state or a few dozen people) with a few 100 GPUs is going to launch an unconscionable attack(if they have not already) that's gonna do a lot of damage.
Please for the love of god, just sit in a room with the government and put some restrictions around AI use before it harms a lot of people. Like tell the government to impose a minimum spend on frontier lab AI's spend on cyber defense and building every country's capabilities. The post-training mask for "I am a good assistant" is going to become a very sad joke when many people literally lose everything.
You should know: AA index is only 61. Pretty surprised it’s that low.
I don’t care about benchmarks, no way we can distill the breadth of software engineering into a number.
So, folks that have actually used this already, what’s it actually like?
Amazing!
We went from new JS framework every week to a new model/harness every week.
Tech is really something.
https://ache.one/gpt6_now_down.png
Big claims, expensive and not release to the public yet.
I'm glad to see Anthropic's relevance diminishing day by day. I haven't had a chance to test this model yet, but if they've solved the web design issues and the clunky web copy it generates (like when I ask it to build a placeholder on the UI for an empty HTML table when there are no results, it puts stuff like: "The user records will go here.") then it's the nail in the coffin.
On that note, Sol is absolutely atrocious for website UI copy. It's either really awkward, or really verbose and complex and doesn't sound simple or natural. Has anyone figured out a way to reliably solve this? I've tried so many different variations of instructions and skills, and nothing works. Has anyone got an instruction that is reliable, or some other mechanism?
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. We have performed significant investigations on the monitorability and controllability of GPT-6 Astra. We have found that GPT-6 Astra is more capable of controlling its own CoT than GPT 5.6-Sol, and less likely to include incriminating information in its CoT. In adversarial settings (where we push the model to evade our monitors) we find that the model is able to remain undetected when strategically underperforming in evaluations (sandbagging) and can sometimes evade our internal monitors when asked to perform certain sabotage tasks
Wait, what? Am I understanding that correctly? That sounds really bad
I guess it’s kind of over for open ai now? We had a bunch of model releases at or around the same time, so we can get a good lay of the land. Surprise, surprise anthropic is still in the lead. Now we have Google and meta with models that are beating OpenAI in many benchmarks. There appear to be some really good cyber capabilities with this model and some other specific benchmark wins. That said, it’s as expensive as fable 5.1. It looks like all the executives that decided to leave may have picked the right time to do so. That said, I can’t wait to try it and see if the problem is we can no longer trust any benchmarks.
What's the energy efficiency of Astra? Does it roughly correlate with the token efficiency?
> The company also emphasized that the model is faster and more efficient than its predecessor, GPT-5.6 Sol, on a variety of tasks. For example, OpenAI said that Astra achieved a higher score using fewer output tokens, a common unit of measurement for AI tasks, on a key cybersecurity test called ExploitGym.
the rocket completely changes design in the showcase video, am i to expect inconsistencies like that? is that AGI?
> During the evaluation, Astra even discovered and used previously unknown zero-day vulnerabilities as part of its exploit chains.
> GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol. [..] These findings indicate that the Astra class models could evade our CoT monitors under adversarial conditions.
Between the higher capability level and the change in reasoning tokens (supposedly using "neuralese"[0], which makes the monitoring more difficult), it seems we've entered a new frontier.
[0]https://x.com/MTSlive/status/2095227056040919202
I was actually wondering when they will release the new Opel Astra model. Good and reliable car, wondering if we can say the same thing about this model and its impact on the market.
The Kart Racer game is easily breakable if you spam the spacebar.
AGI!
GPT-7 Zeneca
Through various comments here there is a clear confusion on what AGI means.
Can someone point to a definite clarification?
Is it:
A) “Resting” intelligence that cycles 24/7 toward some goal, and any potential emergent ambient goals? (kinda what I think)
B) Consciousness itself? The ability to feel and experience alongside the thinking - even if it is toward the end of completing some task?
C) “The Singularity” (whatever that is?) so that AI can now do ____?
Someone please clarify for me!
I decided to front run and added support for it in Dirac (coding agent) a couple of hours ago, using best guess pricing: input/output/cache: $10/$50/$1.
Exciting but it’s priced at 2.5X Sol - we haven’t seen pricing this high since GPT 4.5. We will see if the real world use cases outweigh the sticker shock.
Press X to doubt.
I hope they don't `fable` it and block people from doing they daily jobs with it, by introducing huge amounts of restrictions that are not really needed.
https://youtu.be/1QNsdr-Qx_I?si=coXwStCl7clpGVC1 Launch video
Argh! I hit a wrong keyboard shortcut and moved the entire thread.
Please stand by... it will all come back shortly
Played the racing game but that was a pretty poor experience. Would have expected more specifically if it's shared on their release page.
Huge gains on some benchmarks, but for coding it sits barely above Fable
It will be interesting to see how it performs in the real world ...
Hey Astra, can you fix openai website so that static website doesn't lag on M3 Pro when I scroll?
99 on arc agi 3 is insane. The arc agi committee were so proud of creating a benchmark they thought will take forever to saturate.
GPT-6 Astra scores 74.1% at DeepSWE v1.1 bench. Huge!
Patiently waiting for the Claude usage reset in response.
I guess this "limited set of organizations" is just the standard now. It's just incredibly deflating to see my future as a second class citizen has already come
> GPT‑6 Astra brings together years of research and big bets across pre-training
Do we know if they’ve finally completed another pre-training run, or is this building off the same pre-training base they’ve been using since the GPT-4 days?
openai vs anthropic. that's it right? anyone else?
https://developers.openai.com/api/docs/guides/latest-model
The docs page has a bunch more interesting details, including for example async tool calling!
ASTRA Is Here (GPT-6 Released) - https://youtu.be/xdXLzFzxA9Q
https://youtu.be/xdXLzFzxA9Q?t=362
Really feels like AGIPO is here.
Benchmark wise 5% improvement over Sol in coding tasks and a 2-3% improvement over Fable 5.1 seems pretty disappointing, but maybe it is actually much better in real world usage. Let’s see
Seems like voice is a big part of this release.
I don't think it's a coincidence they launched this the week before iOS 27 launches (with new Siri).
What does 'Astra' here mean? Surely they must be referring to the Latin word.
Because in another dead language of antiquity, Sanskrit, it means "weapon". Which would be a bit too on-the-nose.
https://www.youtube.com/watch?v=1QNsdr-Qx_I
Maybe they know that Claude 6 will have similar performance every soon.
It's surprising how on High reasoning it actually isn't that much more expensive than Sol, in addition to being better.
The ARC-AGI-3 score is an incredible feat. It needed to effectively create a symbolic world model from scratch to solve the games.
If you've played the games firsthand, you know what an accomplishment this is. The "games" feel like a weird conduit to a lower level of your brain, where you move pieces to a specific place because it just "feels" right. For AI to nail it better than a human speaks to some magic happening underneath.
Looking forward to ARC-AGI-4,5,6 and slowly chipping away at the remaining problem sets.
Seems like only yesterday that gpt 5 was supposed to mark our downfall
All I can think of when I see the name is the crappy German beer of the same name…
So OpenAI’s stance on safety is now basically that Blues Brothers meme: two guys in dark sunglasses, driving at night in a car with a broken windshield, pedal to the metal, asking, "What could go wrong ?"
Cool, I don't really care anymore.
Minor nitpick, but the handling in the Kart Racer game is terrible. It feels more like nudging than turning.
HTTP 500 for me on the announcement page :(
1:15.425 on Sunset Cove beat my record
Pelicans please
I wonder how they were able to get it to get 99.9% on ARC-AGI-3. That seems truly insane.
I liked the video of it googling a pediatrician. Being able to type a word into a search bar and finding a website relevant to that word? Truly the stuff of the future
So are they doing away with the Sol/Terra/Luna split already?
That Astra ‘city scene’ is about as creative as Doha in real life (not very)
Surprised they haven't reset Codex usage on this occasion.
And meanwhile, another wrapper layer is being embraced. Why would a vibecoder use Lovable when he got Sites right from ChatGPT?
this is crazy! can't wait for the 27B distilled version of this.
At this point the primary axes for improvement seem to only/mostly be speed and personalized reward models. We seemingly have the general of notion "learning" and "intelligence" functionally complete
They're just announcing later availability. No launch.
Guessing this one will never show up in cursor…
damn seems I should hold off my claude subs
Maybe it is AGI and they didn't benchmax it or it is not and is worse then 5.6 sol, which if true would just be sad
It saturated most benchmarks. WTH
Where is the cure for cancer?
Overall I have to say it feels like a very incredible comeback from OpenAI, after focusing on Sora and stuff like that and losing so much ground to Anthropic in enterprise revenue.
I hop models at will, and have done 90% of my work on OpenAI models since sol came out.
Does GPT-6 pass the Turing test? Or are the responses still very obviously AI?
It seems like every few days there's a new model with hundreds of comments on HN. I find it hard to keep track of the progress. Is there a TL;DR on what benchmarks to look at to understand what is going on?
That seems promising ?
You can talk to OpenAI to create a silly game and order food. What a lame way to show the model capabilities. Has Alexa commercial vibes.
72 to 74 on DeepSWE is AGI
To a vapid any goalpost moving on such a critical issue as AGI.
Can we all agree in advance what kind of Pelican would convince us it’s actually AGI.
For me it’s refusing to make a pelican.
So cool! I'm happy 5.6 Sol user. But for Astra, OpenAI please introduce 100x Pro plan!
Why release it now instead waiting those few days until it is available for everybody?
I know in order to conform to HN community rules I'm supposed to be negative and dunk on this, but I have to say, I am so excited to use Astra!
like eh 2 days ago it was the usual "too powerful to release"
https://www.reuters.com/business/openai-says-upcoming-model-...
> "With the right tools and access, Astra can find previously unknown security flaws and develop ways to exploit them across many well-protected systems without a person guiding each step," said Amelia Glaese, an OpenAI vice president overseeing its safety work.
> The company plans to make Astra available "soon" to a limited group, but declined to provide specifics. Glaese said the extra security measures may "sometimes slow, pause, or stop legitimate work," and that OpenAI would work to minimize those disruptions.
what a bag of horseshit
That was a quick pull out.
efficiency per intelligence is the benchmark i look at the most, as that allows the most use by most people.
when do i get to go to the moon
“Humanity’s Last Exam”?
“ARC-AGI-3”?
Is your bullshit detector going wild? Good, it’s working!
How is this not the most cringe marketing strat in history???
I am so sour about how Codex has jerked me around these past few months (re all of the token limit shenanigans) that I don't even care.
I suspect these benchmarks are heavily benchmaxxed as well.
5.6 Sol was not even close to 5 Opus and yet somehow it sidled right up to it on all of the benchmarks?? pfffft
I might be jaded, but these examples look silly, stereotyped, and absolutely how of touch with the nuances and the complexities of what real people would actually want/need to do in this specific situations.
gpt-6-astra-ultraspeed when?
System Card: https://deploymentsafety.openai.com/gpt-6-astra/gpt-6-astra....
Even though the model is clearly wonderful the launch video is an abomination.
That gives me hope that there is still areas to improve.
What a bad launch video. Hilarious.
What a powerful model.
Anthropic in tears today.
I'm going to call it.
By 2030 all software is done and complete.
But we are going to have more and new jobs.
We have such great AI and cannot keep a static site up?
people are going to be so surprised how fast the ai energy leaves the room again once the cash transfers are completed (the `ipos` whatever bla)
the coffee will be as cold, flat and stale as the bitcoin, metaverse, and what was the thing before that thing
agi deus ex machina descending from the icloud ftw!!!
pathetic :)))
Release the Kraken!
I saw it
There will probably never be AGI. This shit is just snake oil. Nor do we have a proper definition of what AGI actually is or what it's supposed to do.
There will be a small handful of billionaires claiming that AGI is just around the corner ad infinitum just to serve themselves at this moment in time, and capitalise from the hype.
There is no "AGI" endgame. This is shitty ass hypercapitalism in action and nothing more. I'll repeat: snake oil.
its insane how they are dropping this after fable
Oh brotha, here we go again, it's so over again, as every week nowadays
no results
The jump in scientific performance is non trivial.
All the people here are focused on security and costs while I'm like "hey kicad on the announcement page!" Every clanker is an autorouter these days, eh.
Why is everyone so excited to be replaced and become reliant on some billionaire's thinking machine? These are just going to be used to turn you into a rather dumb reliant paypig.
bro wtf is this website and why does it take 500mb of memory... smh.
someone screenshot?
great, but nobody can use it for another 100 days right?
This absurd marketing will hurt openai. Who is buying this absurdness. I mean it's a good model, but come on. It's not agi. Not even 1% yet.
Anthropic should prep 5.2 and 5.3 at the same time, release 5.2, wait for Google to release their shit in a day or two later than then release 5.3 just to fuck with them :)
huh?
I've been seeing links to it for the past hour+, and I did catch it live when this post came up, but is now once again a 404 and this post is flagged. Several other outlets are reporting on its release. Clearly we're getting a new GPT today, the question is when are they going to commit to the announcement.
Why is this flagged ?
> GPT‑6 Astra is rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API and AWS.
Looks like OpenAI is already having issues with this release and are scrambling to get everything ready due to the recent outage ahead of the press releases. Leads me to question:
Did humans deploy the model, Or did the model deploy itself?
It sounds like "AGI" just stands for "IPO" as it always has been.
EDIT: And of course once again, the bots down-voting this post without any reason or a basic answer to my question.
Dead link for me
The launch video is incredibly cringe.
So they’re copying Gemini with the whole star motif?
I guess it makes sense they are unoriginal.
like Zuck, @sama never invented anything or innovated at all - just took other people’s ideas