“Artificial Analysis”, the most closely watched independent evaluation platform in the artificial intelligence sector, has updated its “Intelligence Index” scoring system three times in less than a week. In the v4.3 version released on September 7, OpenAI’s “GPT-6 Astra” model and Anthropic’s “Claude Fable 5.1” model tied with 53 points each.
“Elchi” reports that “Artificial Analysis” officials stated that versions v4.1.1, v4.2, and v4.3 were released as part of an accelerated process leading to a more comprehensive “v5” update. They noted that due to the rapid pace of progress in the industry, it was not possible to delay the update.
What changed during the three updates in one week?
At the beginning of the update process—when “GPT-6 Astra” was first introduced—Astra had scored 61 points in the v4.1.1 version, while the “Fable 5.1” model had 66 points.
In the v4.2 version released four days later, the platform added new evaluation criteria and simultaneously removed the “GPQA Diamond” test, which had already reached a saturation point. At this stage, the gap narrowed: “Fable 5.1” scored 57 points, and Astra scored 55 points.
Three days later, with the release of the v4.3 version and the integration of “Terminal-Bench” and “AutomationBench-AA” tests into the system, both models tied with 53 points.
Şayəstə