20 points cdnsteve 2 hours ago 6 comments
varispeed 56 minutes ago | parent
These tests are pointless when often models get nerfed few days after release. Astra today is way dumber than just few days ago.
cbg0 39 minutes ago | parent
Not Astra but Sol hasn't been nerfed: https://marginlab.ai/trackers/codex/
d_tr 39 minutes ago | parent
How and why do they get nerfed? To save money?
kzrdude 37 minutes ago | parent
Is that something we have credible evidence for? Do they serve a better model when artificialanalysis (the benchmark site) is making the requests, and so on?
forgot-my-pw 19 minutes ago | parent
Terminal Bench 2/2.1 appears to be almost solved, so we probably shouldn't look too hard on that? Their Terminal Bench 4 score is soso.
Still quite impressive though.
walrus01 51 seconds ago | parent
Not really news, terminal bench 4 is the new metric. It's only a few points ahead in terminal bench 4 of some open weight models you can run on a 256GB system.