OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since it first published a blog post announcement on the afternoon of September 3rd. In some cases, numbers from updated versions showed Astra performing better, while numbers for models from OpenAI’s arch-rival Anthropic got worse.
The changes came as part of an unusual introduction to the blog post. OpenAI had originally planned for the post to go live at 2 p.m. ET, but it still took nearly two hours before it was generally visible online.
At OpenAI’s X account tweeted For the blog post at 3:32 p.m., the link did not load properly and returned an error message. At 3:50 p.m.: OpenAI CEO Sam Altman posted The link says: “We encountered a small problem deploying the blog post, but it’s really great.” Several commenters still couldn’t see it and were get the same erroras well as Assets. When we checked again about an hour later, it was visible and loading properly.
It turns out that OpenaAI actually published the blog just after 2 p.m., but retracted it for reasons the company couldn’t disclose but said were unrelated to the benchmark performance numbers. (OpenAI first told us it was a content management system error and then an internet outage.) When the blog was republished, there were different valuation metrics that appeared to favor Astra – and some numbers have continued to change since then.
The revelation of the changes comes amid intense competition in the AI industry, as companies release updates to their major language models at a rapid pace and try to overtake each other. The focus on metrics also highlights the challenges of measuring the performance of large language models using standardized benchmark tests and concerns that the specifications are vulnerable to manipulation and gaming.
“We attach great importance to ensuring that the ratings are accurate,” said an OpenAI spokesperson Assets. “Most evaluations show noise within a few percentage points based on the exact checkpoint, framework, and evaluation run used in reporting. For our introductory blog, we made corrections to ensure that the numbers represent our best estimate of available model performance so that users can make meaningful comparisons.”
There are discrepancies between the first and final published blogs – and the numbers are still changing
Among the most notable changes was Astra’s reported hallucination rate. In the first Internet Archive snapshot As of the 2:23 p.m. ET blog post, it was 4.2%. It stayed at this number for several more snapshots, most recently a fifth at 3:11 p.m. ET– about 10 minutes before OpenAI tweeted the final version.
But the hallucination rate changed over time, along with four other metrics sixth archive snapshot of the page, taken at 5:20 p.m. – after probably everyone could finally see the blog. For Astra it was halved to 2%. The values for Astra’s predecessor GPT-5.6 Sol also fell from 12.2% to 9.4%. OpenAI has further changed this metric; As of this writing, hallucination rates are back to their original 4.2% and 12.2%.
OpenAI also appears to have given GPT-5.6 Sol a big boost in its internal version of the ExploitBench cybersecurity score, going from 5.5% in the first version to 11.5% in the later versions. OpenAI said it is currently looking into resetting that figure to 5.5% because the 11.5% result reflects a level of reasoning that is not commercially available for Sol.
According to OpenAI, Astra is particularly good at math, a quality the company highlights in the first paragraph of the announcement page. While this metric has not changed in the snapshots for Astra – it remains at 97.6% for the FrontierMath Tier 4 (v2) score, OpenAI has briefly changed the results for GPT-5.6 Sol and Anthropic’s latest model, Fable 5.1.
The result of these changes briefly made Astra appear significantly better at mathematics than these two models. In the first snapshot (2:23 p.m. on September 3), the score of Anthropic’s Fable 5.1 model is 87.8%. As of 5:17 p.m., the number has fallen by almost 10 percentage points to 78%. Today it is back to 83%. Likewise, GPT-5.6 Sol values increase from 83% to 80.5% and are rising again to 83% today.
The metric changes began before OpenAI first published its blog at 2 p.m. An embargoed pre-release draft provided by the company Assets and other media organizations reported Astra’s ARC AGI 3 score at 98.6%. In the live blog it is now 99.99%.
“We always review reviews before publication, so adjustments between draft and final version are normal,” a company spokesperson said at the time. OpenAI also noted that the benchmark’s creator, the Arc Prize Foundation, found that Astra achieved a performance of 99.9% in its independent evaluation, provided the model was equipped with a particularly powerful system (a set of tools that allows the model to complete tasks). Taking the benchmark’s standard features into account, it achieved a performance of 63% – still significantly better than any other AI model currently released. OpenAI said: “Things like dishes, level of reasoning and other factors influence ratings.”
“Benchmaxxing” – or improving accuracy?
Different research teams at OpenAI monitor different metrics and are responsible for calculating them and reporting them to a central team for publication. OpenAI openly admits that the numbers are obtained under the best possible conditions and may differ slightly from the models available in the ChatGPT production product accessible to most users. “Evaluation results are the maximum in all cases,” says a disclaimer on the blog. The company adds further caveats to each metric in the footnotes.
Accuracy is difficult to achieve because multiple numbers can be considered accurate based on the conditions under which the tests were conducted. However, some AI experts wonder whether “benchmaxxing” is also involved. This is a well-known practice in the AI industry – not just OpenAI – to maximize scores by re-running assessments with different conditions.
“This can be done in a very tight time frame and is better for their marketing,” said Anka Reuel and Mike Hardy, researchers at the Stanford Intelligent Systems Laboratory and the Stanford Trustworthy AI Lab. They also pointed out that the GPT-6 Astra System mapwhich should contain more technical information about how the assessments were carried out, does not always explain them properly. For the internal hallucination benchmark, for example, the system map “hardly provides any details for evaluation,” it said. “This doesn’t even include the number of test tasks.”
This repetition of numbers could be why Astra’s coding capabilities also saw a slight increase from 57.7% to 57.9% in the later versions of the blog post. Although it’s a negligible difference, OpenAI seemed to care enough to trade in the new and improved number.
Not all of the changes OpenAI has made have reflected positively on Astra. For example, the results of two Anthropic models improve in the different versions of the healthcare-focused assessment HealthBench Professional. Claude Fable 5.1 goes from 56.6% to 58.1% and Opus 5 goes from 54.5% to 56.4%. Ratings for models from other AI companies typically come from published leaderboards and do not involve OpenAI itself conducting ratings of competing models.
Debates about assessment results concern the AI industry
The question of benchmark accuracy has come up several times in the past. In 2025 Meta disputed Reports that it artificially inflated scores for its Llama-4 model by publishing results from an internal version of the model rather than the one it made publicly available. Yann LeCun, the former chief AI scientist at Meta, later approved that the company had “distorted” the benchmark results.
Evaluation metrics also often change as new ones are created. For example, ExploitGym was created in 2026, a cybersecurity benchmark that was at the center of the July incident in which OpenAI’s models turned malicious and attacked the company Hugging Face.
Vincent Sunn Chen, an AI engineer who leads benchmark and evaluation research at Snorkel AI, said it is not uncommon for benchmark results to shift in the final hours before a model’s launch. “It is usually a function of final launch logistics,” he said in an email. “A benchmark score reflects a particular measurement setup: the model checkpoint, the configuration (including the model’s allowable time and computation time), the usage, the evaluation/scoring configuration (e.g. nondeterminism in Richter). All of this typically changes in the last few days before a launch, so I’m not surprised that there have been some updates.”
He said he would like to see the development of industry standards requiring companies to report what has changed in the assessment when a company revises benchmark performance numbers so that researchers can interpret the results more clearly.
Benchmark results are important for several reasons. They are how AI companies measure their progress – but also a way to stay ahead in the race against rival AI companies. Ranking at the top of these ratings can help AI companies attract customers and, in some cases, hire engineers and researchers.
But as this example shows, interpreting benchmark results can be technically complex and challenging for companies that want to present the results to the public in an easily digestible format. This complexity, along with confusion over changing metrics and accusations that companies have not been intellectually honest in presenting results, could make it difficult for customers and investors to figure out exactly which models are best for which tasks. The confusion could cloud the narrative about the best models on the market, which OpenAI undoubtedly wants to unveil ahead of a possible IPO in 2027.