FOUNDING OFFER · 3 MONTHS
FOR $45 $17.76
CEO TIMES
JOIN NOW
CEO Times
Sign Up
Markets & FinanceBusiness & CorporatePoliticsThe WorldOpinion
NOW
U.S. National Debt Crosses $40 Trillion as Boomer-Era Policies Drive 81% of Future Spending GrowthBillionaire Igor Tulchinsky Donates £5M to British Museum's Bayeux Tapestry Show — the Biggest European Exhibition of 2026Oil Hits $99.85 a Barrel — Up More Than $33 in a YearMystery Nonprofit Drops $2M Bitcoin Ad Blitz in the Wall Street Journal — and Nobody Will Say Who's PayingHunter Biden Launches $LAPTOP Meme Coin — 1 Billion Tokens, 30% Kept by FoundersAdaptability Over Forecasting: Top Executives Declare Certainty a Dead StrategyGavekal's Gave: Chinese Bonds Offer Safe Haven as U.S. Debt Hits $40 TrillionCanada Reroutes $10B in Oil East as U.S. Tariffs Hit 50%Macau Bets $16 Billion to Reinvent Itself as a Business City by 2030Peru's Inflation-Targeting Model Cannot Fix Venezuela — Here's WhyU.S. National Debt Crosses $40 Trillion as Boomer-Era Policies Drive 81% of Future Spending GrowthBillionaire Igor Tulchinsky Donates £5M to British Museum's Bayeux Tapestry Show — the Biggest European Exhibition of 2026Oil Hits $99.85 a Barrel — Up More Than $33 in a YearMystery Nonprofit Drops $2M Bitcoin Ad Blitz in the Wall Street Journal — and Nobody Will Say Who's PayingHunter Biden Launches $LAPTOP Meme Coin — 1 Billion Tokens, 30% Kept by FoundersAdaptability Over Forecasting: Top Executives Declare Certainty a Dead StrategyGavekal's Gave: Chinese Bonds Offer Safe Haven as U.S. Debt Hits $40 TrillionCanada Reroutes $10B in Oil East as U.S. Tariffs Hit 50%Macau Bets $16 Billion to Reinvent Itself as a Business City by 2030Peru's Inflation-Targeting Model Cannot Fix Venezuela — Here's Why
CEO Times
Sections
The outlet
Business & Corporate

OpenAI Changed Astra's Benchmark Numbers Six Times After Publishing Launch Blog

Archived snapshots show GPT-6 Astra's hallucination rate cut in half and rival Anthropic's math score docked nearly 10 points after the model's Sept. 3 debut - the kind of self-reported figures investors now use to price the entire AI race.
Imagen ilustrativa
Saturday, September 5, 2026

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model multiple times since first publishing a blog post announcement on Sept. 3, according to Fortune, which tracked the changes through internet archive snapshots.

The rollout itself was unusual. OpenAI had planned to publish the post at 2 p.m. ET. It went live shortly afterward but was retracted for reasons the company told Fortune it could not disclose, saying the retraction was unrelated to the benchmark figures. When CEO Sam Altman posted the link again at 3:50 p.m., writing 'We hit a little snag getting the blog post deployed, but it is really great,' many users, including Fortune, still could not load it. It became visible roughly an hour later.

The numbers changed along the way. Astra's reported hallucination rate stood at 4.2% through five archived snapshots, the last taken at 3:11 p.m. ET. In a sixth snapshot at 5:20 p.m., after the blog was widely viewable, the figure was halved to 2%. The rate for predecessor GPT-5.6 Sol dropped in tandem, from 12.2% to 9.4%. As of publication, both figures are back to their original levels: 4.2% and 12.2%.

On the ExploitBench cybersecurity evaluation, Sol's internal score rose from 5.5% to 11.5% between versions. OpenAI told Fortune it is now investigating reverting the number to 5.5%, saying the higher figure reflects a reasoning level not commercially available for Sol.

Math scores moved too. Astra's own FrontierMath Tier 4 (v2) figure held steady at 97.6%, but Anthropic's Fable 5.1 model dropped from 87.8% in the first snapshot to 78% by 5:17 p.m., before settling at 83% today. Sol's own math score fell from 83% to 80.5% and back to 83%. On the ARC-AGI-3 evaluation, an embargoed pre-publication draft given to Fortune listed Astra at 98.6%; the live blog now shows 99.99%. The benchmark's creator, the Arc Prize Foundation, found Astra scored 99.9% independently when given a powerful tool harness, and 63% under the benchmark's standard harness - still, OpenAI noted, ahead of any other publicly released model.

An OpenAI spokesperson told Fortune: 'We care deeply about getting evaluations right. Most evaluations have noise within a few percentage points based on the exact checkpoint, scaffold, and evaluation run used in reporting. For our launch blog, we made fixes to ensure the numbers represent our best estimate of available model performance, so that users can make meaningful comparisons.'

The numbers come first, and in this case they moved in OpenAI's favor while a direct rival's moved down - then some moved back. Enterprise buyers and capital markets pouring billions into AI infrastructure treat these self-reported figures as the closest thing the industry has to audited performance data. When a hallucination rate swings by half within hours of publication, without disclosure at the time, that is not statistical noise; it is a governance failure dressed up as one.

Free markets price risk on the assumption that disclosed numbers mean something. A chatbot maker's benchmark page has become, in practice, an unregulated corporate filing that moves stock narratives and procurement decisions. Power leaves a paper trail, and in this case the trail runs through six archived snapshots, not a press release. Investors deploying capital into the AI race should treat vendor-reported benchmarks with the same skepticism they once reserved for non-GAAP earnings adjustments - and demand the receipts.

More from Business & Corporate