140 points meetpateltech 1 hour ago 75 comments
ls1911 1 hour ago | parent
simianwords 57 minutes ago | parent
The personality is bland and it doesn’t work nearly as hard or even tries to help.
artemonster 44 minutes ago | parent
slowin 43 minutes ago | parent
moojacob 53 minutes ago | parent
Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.
However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.
My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.
jasonjmcghee 42 minutes ago | parent
That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.
vessenes 18 minutes ago | parent
dumberquestions 34 minutes ago | parent
user43928 15 minutes ago | parent
How representative that is of real world usage, I don't know.
In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.
Lucasoato 15 minutes ago | parent
I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?
Aperocky 12 minutes ago | parent
When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.
grababner 7 minutes ago | parent
atniomn 14 minutes ago | parent
kristofferR 49 minutes ago | parent
Iolaum 45 minutes ago | parent
babelfish 44 minutes ago | parent
Jcampuzano2 42 minutes ago | parent
Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.
Jcampuzano2 41 minutes ago | parent
This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.
kristofferR 20 minutes ago | parent
If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.
andsoitis 17 minutes ago | parent
kristofferR 15 minutes ago | parent
user43928 8 minutes ago | parent
That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.
scottyah 35 minutes ago | parent
kristofferR 17 minutes ago | parent
sidgtm 30 minutes ago | parent
vessenes 20 minutes ago | parent
Tsarp 17 minutes ago | parent
rvz 12 minutes ago | parent
At this point you might as well ask an AI model to generate audio waveforms from text and judge it as an audio model or ask a model specifically designed to generate SVGs [0] to generate videos.
jcims 6 minutes ago | parent
Saline9515 17 minutes ago | parent
It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.
andsoitis 16 minutes ago | parent
AM1010101 14 minutes ago | parent
maz1b 14 minutes ago | parent
6thbit 12 minutes ago | parent
meerita 11 minutes ago | parent
zug_zug 10 minutes ago | parent
I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?
grim_io 4 minutes ago | parent