SpaceXAI's Grok 4.6 Scores 61 on the Artificial Analysis Intelligence Index (artificialanalysis.ai)

67 points by wertyk an hour ago

25 comments:

by sidcool 3 minutes ago

Grok is not the best model around, but it's decent. It gets the basic job done at a low price. I don't think it can advance frontier Math, yet.

by satvikpendem 34 minutes ago

Cursor, since Grok 4.5, has had an incredible deal for frontier level models, their subscription now goes way further than OpenAI or Anthropic. Even on their lower tier plans you can use a lot tokens on their of their first party models (Grok and Composer) and not really run out comparatively. Combine them with an orchestrator and implementor type setup and it goes even further.

by aliljet 31 minutes ago

Can you explain what you mean? These days courtesy of an addictive reset game OpenAI is playing, I can't find anything with frontier intelligence that's more cost efficient...

by jesse_dot_id 28 minutes ago

Goes even further to exfiltrate your data, yeah.

by greenavocado 8 minutes ago

That would be Muse Spark Contributor Tier. 12-21x price reduction at the expense of your digital existence.

by nylonstrung 31 minutes ago

I have never met a single human being who uses Grok for coding

by jm4 20 minutes ago

I have. He was using it due to philosophical reasons the same way many people have philosophical reasons for avoiding it. I don't know how many people are like that, but it's not exactly where you want to position your product if you're a business.

Personally - and I know I'm not alone with this sentiment based on comments I see on this site - I wouldn't touch Grok no matter how good or cheap it is. I don't trust Elon and I don't want to give another dollar to the world's richest person who turns around and uses the money to interfere with elections. The guy I know uses it for essentially the same reason I won't use it.

by busch_j 22 minutes ago

A bunch of SWEs at my work use it as their primary model.

We have Claude, ChatGPT, and Cursor with essentially no cap on spend (top guy is spending over 10K a month on AI at API prices), and he hasn't had his hand slapped.

So it's not like they are using it purely because it's cheaper.

I think people like to use it for its speaking style, pretty solid performance, and its speed.

by visopsys 27 minutes ago

I use. I used to be a Claude user. Since trying Grok 4.5 and especially Grok 4.6, I don't want to go back to Claude any more (I have early access to 4.6).

Grok is 3x+ faster than Claude and I can't tell the diff in engineering work quality. As an engineer, speed is important to me.

by bigyabai 25 minutes ago

For $30/month, I'd expect it to have higher usage limits than Claude Code and Codex.

by visopsys 17 minutes ago

It really does. I felt like I could have spent $1000+ api token on claude for the amount of work on my $30 grok subscription.

by drewnick 10 minutes ago

An hour in, I've been running four terminals full bore on my $20/mo Grok sub and I'm at 9% for the week. Codex or Claude would easily have hit 5-hour or weekly limits.

by peder 16 minutes ago

Does it matter? Why turn it into a popularity contest?

by LeBit 22 minutes ago

I refuse to use that product because of the parent company.

by Recurecur 4 minutes ago

I wonder why folks feel the need to virtue signal like this…?

As an aside, do you use Apple equipment despite Apple’s use of Chinese slave labor?

by Recurecur 8 minutes ago

Hi! Grok’s worked quite well for my use cases…

It’s also a great deal!

by kvirani 28 minutes ago

Folks working in US govt tend to, based on convos I've had with one such person.

by sidcool 30 minutes ago

Hello. Nice to meet you.

by supriyo-biswas 29 minutes ago

I'm only being forced to use it at $WORK since some people overran their Cursor bill, so everyone gets Cursor Auto enabled by default which routes to Grok 4.5.

by DetroitThrow 27 minutes ago

I've tried it on my "let's run every model in parallel and see which finds more edge cases" type of tasks, and Grok 4.5 was really behind Opus/ChatGPT but ahead of Gemini - despite having a strong showing on benchmarks.

That makes me really skeptical of it being GPT5.6-tier, much less Fable-tier, based on some of these benchmarks alone. But I'll test here shortly.

by pzo 24 minutes ago

Seems the cache read pricing almost doubled from $0.30 in Grok 4.5 to $0.50 in Grok 4.6.

In my experience in heavy coding sessions most pricing is just cache read and cache write like 80% of my token bill.

by thiago_fm 34 minutes ago

I often wonder if there's a chance, even if minimal... that they stole the weights of the Anthropic models they run on their datacenter... or are actively destillating it.

by connicpu 31 minutes ago

I think the more likely explanation is that the Cursor data they effectively acquired for $10B was extremely valuable for their training when combined with the insane number of GB300s xAI has for training.

by winstonp 23 minutes ago

Cursor was 60B. The 10B number was the breakup fee if the deal fell through.

by qudat 23 minutes ago

> ... or are actively destillating it.

I just assumed every model manufacturer is distilling from the frontier models. If they aren't they are definitely trying to do it.

Data from: Hacker News, provided by Hacker News (unofficial) API