GPT-6 Astra vs. Claude Fable 5.1: Who Actually Wins, and Where the Hype Gets Ahead of the Data

Image by BoliviaInteligente on UnSplash

GPT-6 Astra made its very noticeable debut on the 3rd of September, with Sam Altman posting an apology for a messy rollout. OpenAI definitely did not bother with trying to be modest either. The statements from Open AI predictably call it the most intelligent and aligned model in the company’s history. The president of OpenAI, Greg Brockman, told reporters that it’s “not unreasonable” to assume that we have entered the AGI era. These are bold words that pique the interest of skeptics like me. So I checked the scoreboard just to be sure.

Scoreboards have a habit of telling messier and more fascinating stories than press releases, and this one was no exception. To be fair, OpenAI tooting its own horn is justifiable. Astra does leave other models in the dust in a handful of categories, some of which are refreshingly weird. But Fable 5.1 wins when it comes to the independent, vendor-neutral ranking of general intelligence. Let’s talk more about it.

Where Astra Stands Out: 3D and CAD

Astra handles professional design software the same way that a human engineer or digital artist would, it doesn’t just generate an unchangeable, flat 3D shape.  It also has the ability to self-correct and run quality control tests. If Astra designs a 3D object and flags some parts as “weird”, it will trigger a test render, identify its own mistake, and rewrite the software code to implement the corrections.

BenchCAD: a benchmark that has models reconstructing CAD programs from rendered 3D views. Upload a picture, and Astra can reverse-engineer a shape from the picture. Astra scores an impressive 95.9% versus Fable 5.1’s roughly 84-85% according to Vellum, which is a significant double digit gap.

A point worth noting is that OpenAI does note that the Claude comparison runs used modified evaluation settings, so this particular gap should be treated as directionally true rather than as gospel.

Astra also posted a big jump on AutomationBench (41.4% vs. 31.4% for Fable 5.1) and ScreenSpot-Pro, a test of clicking the exact right pixel in a cluttered UI, where it hit 92.7% against Fable 5’s 87.3% (DataCamp). 

In layman’s terms: if your job involves wrangling desktop software, CAD files, or UIs, Astra would make the perfect assistant.

The Coding Comparison: Closer Than the Highlight Reel Suggests

Fortunately (or unfortunately depending on where you stand), there is no SWE-bench bloodbath. This cycle, the benchmark in use is DeepSWE v1.1, a 113-task agentic coding test, in addition to Terminal-Bench 4.0 and an aggregate Coding Agent Index. On DeepSWE, Astra  scores 74.1% compared to Fable’s 67.4% in OpenAI’s own table. 

A side note worth adding is that independent trackers put Claude Opus 5 73.7% and Muse Spark 1.3 from Meta at 75.4% for the same test. So we can conclude that Astra’s lead in general in this aspect is more of a rounding error, and not a definite knockout.

Where Astra clearly pulls ahead is Terminal-Bench 4.0, which is the test for the longer and messier multi-step terminal work. Astra scored 58% compared to Fable 5.1’s 55.8%. Of course, the win is modest, but Astra does legitimately seem to be better when it comes to not randomly losing its train of thought part-way through a particularly complex shell session.

Image by Ilya Pavlov on UnSplash

But let’s zoom out a bit and look at Artificial Analysis’ Independent Coding Agent Index which combines a number of these tests, and Fable 5.1 pulls ahead and leads 70 to 67.

Where Astra gains full bragging rights is the efficiency, not the raw score. On efficiency, Astra easily matches Fable 5.1’s performance in coding at less than half of the token cost. A pitch that leads with cheaper, not smarter, is still a decent pitch, but it’s just not the one Open AI highlighted or led with.

Artificial Analysis’Intelligence Index (which happens to be the broadest and most cited yardstick in the industry and notably not run by either competitor) has recorded Fable 5.1 beating Astra 65.7 to 61.2. Fable 5.1 also wins Humanity’s Last Exam with a very clear margin at 65.0% to 57.2% according to Analytics Vidhya. In light of all this, the claim about Astra being the “world’s most intelligent model” is simply more about OpenAI’s own tables and less about the claim being something that the entire industry agrees on.

A Note on the Music and Art Claims

There is a likelihood that you may have bumped into some viral posts claiming that Astra aced a “Bach Benchmark,” reconstructed entire cities in Unreal Engine, or built a 3D animated pianist. This sounded phenomenal, so I looked into it. Unfortunately, I could not reliably verify these claims against any actual benchmark organisation or vendor source. 

One of the viral threads had to add a correction noting that a widely shared “Astra gameplay” clip was actually an AI-generated video with watermarks from an unrelated tool, not a real model output. Given that, I’m leaving those claims out rather than repeating unverified social-media demos as fact. I also could not find any published benchmark, from any organization, testing either model on “painting people from pictures” — that appears to be internet speculation rather than a real evaluation, so it’s excluded here too.

The Stuff Only Astra Currently Does

A few genuinely distinct capabilities showed up in the launch data:

  • Cybersecurity, for better and worse.  Astra is the first OpenAI model to cross the “Critical” threshold on OpenAI’s own Preparedness Framework for cyber capability, solving 88% of reverse-engineering tasks on SRE-Bench in a single attempt (OpenAI). That’s a genuine capability leap, and it’s exactly why access to the advanced version is gated behind a vetted program rather than handed out freely.
  • Math that’s actually saturating benchmarks.  97.6% on FrontierMath Tier 4, a test built specifically to resist AI, versus 87.8% for Fable 5.1 (DataCamp).
  • Half the hallucinations.  Astra’s hallucination rate on Artificial Analysis’s knowledge benchmark dropped from 92% to 51% at maximum effort — a real improvement, though a 51% rate is still nothing to hand your homework to unsupervised (Artificial Analysis).

So, Does Astra “Outdo” Fable?

In the categories OpenAI chose to spotlight (CAD reconstruction, terminal work, math, and cybersecurity), the answer is clearly yes. On the benchmark built by neither company to be the fairest overall measure of general intelligence, Fable 5.1 wins, and it isn’t close. Astra’s real headline isn’t “smarter than everything,” it’s “specialized, efficient, and occasionally overpriced for the gains it delivers.”

Worth remembering, too: Astra costs 2.5x more per token than its own predecessor, and independent labs like Artificial Analysis explicitly frame their numbers as reference points, not purchase decisions (Emergent). I would recommend that you select a model based on workload rather than based on a press release

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *