• Not_mikey@lemmy.dbzer0.com
    link
    fedilink
    English
    arrow-up
    1
    arrow-down
    5
    ·
    7 hours ago

    The models are still advancing, especially in math and “reasoning”. For example the frontier math benchmark is showing big strides for the frontier models. A year ago they were only getting <10% of questions, now sol from openai is scoring 89% and they’re reporting the unreleased astra is at 98% .