Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58).
In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.
It also cost more to run the suite than Fable 5.1.
Am I missing something or is this not looking too... stellar?
Well, seems like ECI [0] and the AA index is diverging quite a bit. Benchmarking LLM is tough and I think we are seeing the limitations of current benchmarks and applicability to real tasks.
[0]: https://x.com/EpochAIResearch/status/2095602754282783108
Title: "major gains"
First chart: from score 61 (GPT-5.6 Sol) to drumroll 61 (GPT-6 Astra)
Indeed. Though to be fair it is referring to "Artificial Analysis Coding Agent Index", from 65 to 67.
I think they mean cost per task, where Astra is now on the Pareto frontier.
In the general Intelligence Index it scores exactly equal to Sol (61). In the Agentic Index it scores significantly lower than Sol (51 vs 58). In both it scores lower than Fable 5.1, Opus 5 and even Muse Spark 1.3.
It also cost more to run the suite than Fable 5.1.
Am I missing something or is this not looking too... stellar?
more like 5.7 not 6
5.6.1