OpenAI quietly boosts some of Astra’s evaluation metrics, and continues to change others post-launch

8 hours ago 3

OpenAI has changed several evaluation benchmarks for its GPT-6 Astra model since first publishing a blog post announcement mid-afternoon on Sept. 3. In some cases, the numbers on the updated versions showed Astra performing better, while numbers for models from OpenAI’s arch rival Anthropic got worse.

The changes occurred amid an unusual rollout of the blog post. OpenAI originally planned for the post to go live at 2 p.m. ET, but it took almost another two hours before it was widely viewable online.

When OpenAI’s X account tweeted out the blog post at 3:32 p.m., the link was not loading properly, returning an error message. At 3:50 p.m., OpenAI CEO Sam Altman posted the link, writing, “We hit a little snag getting the blog post deployed, but it is really great.” Multiple commenters were still unable to see it, and were getting the same error, as did Fortune. When we checked back about an hour later, it was visible and loading properly.

It turns out OpenaAI actually published the blog shortly after 2pm but retracted it for reason the company said it could not disclose, but which it said were unrelated to the benchmark performance figures. (Op...

Read Entire Article