30 points | by matt_d 2 hours ago
5 comments
The fact that opus 5 is outperforming fable is odd to me
From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
Sad to see no mention of Gemini ...
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.
The fact that opus 5 is outperforming fable is odd to me
From personal experience, opus 5 feels net inferior to fable on almost every aspect (for coding tasks)
Evals on actual research workflows is the right direction, most agent benches are toy tasks.
I'm glad to see that GPT Sol beats Opus at least in Mathematical Sciences, because that's my need right now, and I much prefer GPT's prose style.
Sad to see no mention of Gemini ...
Damn. These things aren't AGI... but I don't care.
Luna is good enough for me to give a parser spec and have it write one.