119 points theanonymousone 10 hours ago 46 comments
jampekka 8 hours ago | parent
throwa356262 8 hours ago | parent
Human error means this wasn't just stopped together by some bot.
jampekka 7 hours ago | parent
theanonymousone 7 hours ago | parent
MichaelNolan 22 minutes ago | parent
egeres 6 hours ago | parent
SyneRyder 5 hours ago | parent
The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.
seahorseemoji 3 hours ago | parent
I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.
sipjca 1 hour ago | parent
conception 1 hour ago | parent
dom96 6 hours ago | parent
KillSwitch-Bench 1.0
Claude Opus 5 66.9
GPT-6 Astra 57.9
Claude Fable 5.1 46.7
MiMo-V2.6-Pro 38.8
Muse Spark 1.3 36.5
1 - https://bench.killswitch-lang.org/ricardobeat 3 hours ago | parent
drittich 1 hour ago | parent
conception 1 hour ago | parent
gandreani 35 minutes ago | parent
Gareth321 6 hours ago | parent
unsupp0rted 5 hours ago | parent
I've stopped using Astra entirely and remain on Sol orchestrating Luna Xhigh, but it's still not nearly a week's usage for a week's allotment.
And even then, whenever a new model is about to come out, it feels like the model I'm using is being dumbed down substantially.
I have no evidence for this and can have no evidence for this, but I can vote with my wallet regardless.
Gareth321 5 hours ago | parent
Even when I try to stick with Sol X/High, my limits are at best half of what they were before Astra launched, and the intelligence has declined markedly.
I cancelled my $100 plan. This is absolutely absurd and frankly unusable now.
Muromec 4 hours ago | parent
Gareth321 4 hours ago | parent
f6v 3 hours ago | parent
sjbzbeiks 1 hour ago | parent
When I’m doing work on a repo where I’m implementing a standard and the agents have to read the standard to keep from hallucinating my usage skyrockets.
Hell this changes depending on which language I’m working with.
christophilus 2 hours ago | parent
dangoodmanUT 1 hour ago | parent
sauwan 58 minutes ago | parent
gandreani 39 minutes ago | parent
I wonder what dangoodmanUT is using! This is the time to compare!
phoghed 3 hours ago | parent
Serious question: does anyone have evidence of this?
It’s something that’s constantly asserted, and has been since 2023. Every time someone posts a site that tries to track this though, I look at it and it’s just a flat line.
conception 1 hour ago | parent
By and large they don’t. I have seen this drop a few times, eg before fable came out opus dropped a lot probably due to less compute available.
My guess is it’s a combination of getting used to the new cliff models fall off on and forgetting that model performance drops significantly when context fills up.
So new model comes out, people try it and it’s amazing on a task or two. Then they start using it, context window fills up and it gets a lot worse.
loloisi 3 hours ago | parent
But this week they seem to have tweaked the system to a point at which all models (Astra, Sol, Luna) hit rate limits all_the_time without me being anywhere close to the weekly limit.
Early results with MiMo 2.6pro are quite encouraging for anything that's non-UI work so likely switching spend for the time being
ignoramous 5 hours ago | parent
For a model that matches Muse Spark 1.3 in benchmarks, MiMo v2.6 Pro is incredibly cheap, given its cache rates will remain $0.0036 per million.
NortySpock 5 hours ago | parent
imjonse 4 hours ago | parent
drbscl 4 hours ago | parent
Mostly due to lower cost of living; Shenzhen is way cheaper than SV
tensegrist 3 hours ago | parent
segmondy 2 hours ago | parent
deanc 1 hour ago | parent