140 points meetpateltech 1 hour ago 75 comments

ls1911 1 hour ago | parent

after using cursor grok & trae.ai for several months , grok curor is highly superior results to trae.ai

simianwords 57 minutes ago | parent

I guess it’s only my opinion but having used grok for personal chat: it’s by far the worst one amongst Claude, ChatGPT and even Deepseek, Gemini etc.

The personality is bland and it doesn’t work nearly as hard or even tries to help.

artemonster 44 minutes ago | parent

I used openrouter to send same prompt to qwen, derpseek, gemini and grok and found that grok does good research and produces less bullshit, especially when prompted to be critical of an idea

xutopia 7 minutes ago | parent

Ask it to be critical of the birthday photos and see where that gets you.

artemonster 5 minutes ago | parent

can you elaborate?

slowin 43 minutes ago | parent

This has been my experience as well. Grok will end tasks almost immediately and claim "Done!". It's definitely the laziest and most "dishonest" of all the models. The others aren't perfect, but I can't use Grok for any serious coding task.

Capricorn2481 39 minutes ago | parent

> The personality is bland

I don't use Grok, but do you want your LLM to have a personality? "Personality" is exactly what people don't like about Claude.

Razengan 12 minutes ago | parent

I want my sexbot to have a personality

moojacob 53 minutes ago | parent

Apparently Grok 4.7 has 40% more weights than Grok 4.6, but the price ($6 output token, $2 input) is the same.

Given that the decrease in their margin and the fact they delayed the release of Grok 4.7 almost two weeks past the original date, XAI must not have been happy with the results for 4.7. And XAI also waited the day before Opus 5.5 is rumored to launch. I imagine Opus 5.5 will blow Grok 4.7 out of the water benchmark wise.

However, I have become skeptical of benchmarks. Grok 4.5 solved some issues setting up a buildroot system that Fable 5 couldn't do. I find the post cursor groks are phenomenal at frontend web development, though Claude is much better at backend ruby.

My favorite part of the new Groks has been how they speak in plain english. I simply cannot stand Claudish. Or even GPT, which doesn't have Claude's ticks but definitely likes to handwave explaining technical concepts. Still, nothing beats Claude 3.5 and 4 with explaining since it seems all models have regressed. I wonder if Grok 4.7 will also regress with English because of all the RL.

jasonjmcghee 42 minutes ago | parent

For what it's worth - over the last few years or whatever, it seems like Anthropic benchmaxxes the least.

That being said, I currently prefer Sol / Astra to Opus / Fable as I find both to be a better cost payoff to me.

vessenes 18 minutes ago | parent

I was going to say the reverse - claude has been the less satisfying normalized by benchmark for me in the last year. Both astra and fable have their quirks, but I am 90% codex this year up from 10% last year.

dumberquestions 34 minutes ago | parent

Token price doesn't tell you much without knowing token efficiency.

user43928 15 minutes ago | parent

Their leading benchmark with cost per task shows a tough sell compared to Fable 5.1 Low and doesn't reach the performance of Fable 5.1 Medium.

How representative that is of real world usage, I don't know.

In their benchmark GPT 5.6 Sol performs suspiciously poorly compared to the former models.

Lucasoato 15 minutes ago | parent

> I simply cannot stand Claudish

I totally agree, it’s like that as models become more intelligent, they are less understandable by most of people... but aren’t we humans doing the same?

Aperocky 12 minutes ago | parent

The best ideas are usually the simplest to elaborate. If someone comes up with a convoluted scheme that are hard to understand or be adequately explained, it's usually fraud.

When claude speak in convoluted mess, they are often going off on tangents in real work that you asked it to do, too.

grababner 7 minutes ago | parent

If you can't explain it simply, you don't understand it well enough

atniomn 14 minutes ago | parent

I expect the next Anthropic release to finally reduce the prevalence of Claudish

sscaryterry 7 minutes ago | parent

Based on?

moojacob 7 minutes ago | parent

If they fix Claudish, they've earned me back as a max customer!

Fable 5.1 is not there quite there yet.

They need to get that Sonnet 3.5 magic back.

kristofferR 49 minutes ago | parent

What's with the deceptive graph on top? Not including Astra can't have been an oversight, did the model compare poorly to it?

Iolaum 45 minutes ago | parent

I wonder if that means that SpaceX evals show that they consider astra better than fable or that they hate Sam&co so much they don't want to show their stuff.

babelfish 44 minutes ago | parent

this is exactly it.

Jcampuzano2 42 minutes ago | parent

https://openai.com/index/our-decision-on-cursor-following-it...

Its because of this. You can't use Astra in Cursor, and cursorbench uses cursor as the harness. They can't actually benchmark it using their harness hence why its not included.

babelfish 41 minutes ago | parent

They have Astra in other benchmarks lower on the page. They just don't want to show it winning

Jcampuzano2 39 minutes ago | parent

The chart is cursorbench though and they asked about the "deceptive graph"

Jcampuzano2 41 minutes ago | parent

https://openai.com/index/our-decision-on-cursor-following-it...

This explains why. Mentioned in another comment, but cursorbench explicitly tests with Cursor as the harness, and OpenAI doesn't allow them to use Astra in Cursor.

kristofferR 20 minutes ago | parent

That's not accurate. OpenAI doesn't allow Grok to provide Astra to Cursor customers anymore, but it doesn't ban anyone from using Astra via alternative harnesses.

If Cursor wanted to include Astra in CursorBench nothing would stop them, they could easily have spent half an hour vibecoding in OpenAI API key support - if it hadn't been convenient to neglect to do that.

andsoitis 17 minutes ago | parent

Even if they could do that (workaround to include Astra in CursorBench), that has no practical consequences for Cursor users and that's what I as a Cursor user (what I use for dev, though I use ChatGPT for non-dev stuff) care about.

kristofferR 15 minutes ago | parent

It would make the benchmark way better obviously, by showing how their new model compares to their competitors, the whole point of benchmarks and graphs.

user43928 8 minutes ago | parent

> with a proposed shutoff date of November 12, 2026

That said, I don't expect them to benchmark Astra in their Cursor harness given the situation.

scottyah 35 minutes ago | parent

Deceptive? An extremely quick google search would answer your question. OpenAI pulled out of Cursor before they released Astra so it never got that benchmark.

kristofferR 17 minutes ago | parent

Pulled out from letting them resell Astra access, that's not a limitation on running a benchmark.

sidgtm 30 minutes ago | parent

In my experience Grok especially inside Grok build is pretty solid choice, it’s a no nonsense model and stays on its course. Another surface where I truly enjoy the experience of using Grok model is Grok bot

vessenes 20 minutes ago | parent

Nice to see this release cadence increasing and some continued improvement in quality. I am guessing these models are basically still outcomes of the cursor team integrating with the massive amount of compute they now own: I’d imagine we will see significant step up improvements with grok 5 later this year as the team gets more experienced and confident with larger training deployments. Here’s hoping for another competitive frontier model!

Tsarp 17 minutes ago | parent

Waiting on simonw "Generate an SVG of a pelican riding a bicycle " benchmark to judge this model

rvz 12 minutes ago | parent

You mean a performative pseudo-benchmark that tests for nothing.

At this point you might as well ask an AI model to generate audio waveforms from text and judge it as an audio model or ask a model specifically designed to generate SVGs [0] to generate videos.

[0] https://quiver.ai/

jcims 6 minutes ago | parent

We're allowed to have our ceremonies.

Saline9515 17 minutes ago | parent

I tried in Omp (Oh-my-pi), and so far it's really problematic.

It will loop in thinking mode ("Let me implement those fixes: Fix 1, Fix 2, Fix 3 .... Fix 80, Fix 81"), ignore the AGENTS.md instructions, corrupt plan files, etc etc... I have 5.6 Sol as advisor/watchdog, and it blocks every turn, I never saw this. Quite a shame, 4.6 wasn't so bad.

andsoitis 16 minutes ago | parent

Congratulations to the team!

AM1010101 14 minutes ago | parent

Did 4.6 not have an x-high reasoning level? Why are they comparing 4.7 x-high with 4.6 high?

maz1b 14 minutes ago | parent

Either way, the fact that xAI or SpaceXAI or whatever the name is, I can commend the team behind it on their rapid ascent and progress by being close and or on the frontier in several respects.

6thbit 12 minutes ago | parent

( why is the x-axis on the first chart in descending order ? )

meerita 11 minutes ago | parent

Grok it's really expensive. I'm getting really amazing results using DeepSeek 4.1 Flash for fraction of the price.

parineum 7 minutes ago | parent

Brought to you by...

meerita 4 minutes ago | parent

By no one. For the price of 1M token you can get more and with better results with other models.

zug_zug 10 minutes ago | parent

Well I "tried it out" I asked it one question, and it gave no answer and said "Sign up to use more!" I don't think I'll be doing that, no.

I can't think of a single dimension grok is winning on (capability, cost, voice), but want to stay open-minded -- anybody want to vouch for its capabilities in any domain?

grim_io 4 minutes ago | parent

It's probably the most aligned (to a single person) model out there!