20 points cdnsteve 2 hours ago 6 comments

varispeed 56 minutes ago | parent

These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.

cbg0 39 minutes ago | parent

Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/

d_tr 39 minutes ago | parent

How and why do they get nerfed? To save money?

kzrdude 37 minutes ago | parent

Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?

forgot-my-pw 19 minutes ago | parent

Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.

Still quite impressive though.

walrus01 51 seconds ago | parent

Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system.