The models are still advancing, especially in math and “reasoning”. For example the frontier math benchmark is showing big strides for the frontier models. A year ago they were only getting <10% of questions, now sol from openai is scoring 89% and they’re reporting the unreleased astra is at 98% .
The models are still advancing, especially in math and “reasoning”. For example the frontier math benchmark is showing big strides for the frontier models. A year ago they were only getting <10% of questions, now sol from openai is scoring 89% and they’re reporting the unreleased astra is at 98% .