Superconductor
Back to all posts
Benchmark

Grok 4.6 joins the Pantheon

Grok 4.6 is great. It's right up there with Opus 4.8/5, Fable 5, and Kimi K3. It's in the Pantheon.

Sergey Karayev·

Share this:

Another day, another new model release! Between new models (Grok 4.6, Qwen 3.8 Max, Deepseek V4 Pro) and new harnesses (Prime Agent, Deepseek Harness), we’re not hurting for things to benchmark recently.

Today we take a look at Grok 4.6, which we evaluated in two different harnesses: Grok Build and Cursor CLI. Now that Cursor is officially part of SpaceXAI, I guess they will unify the harnesses — and the two CLIs already clash in their alias agent.

As an aside, I’m impressed by how single-mindedly Grok has been pursuing excellence in the coding agent use case, after a late start. I wonder if $60B for Cursor will seem like a no-brainer great deal in the future.

The results

I’ve highlighted Grok 4.6 High and Extra High results on the charts below. We don’t show Grok Build results on the Quality vs Cost chart, because that harness does not report the cost — it uses your monthly plan, and that’s all you know.

Quality vs Time and Quality vs Cost, with Grok 4.6 High and Extra High highlighted

So, two things.

First, Grok 4.6 is great. It’s right up there with Opus 4.8/5, Fable 5, and Kimi K3. It’s in the Pantheon.

Second, the Cursor harness is consistently slower than the Grok harness. There could be two reasons for that, and I have not teased them apart yet. One: the model endpoint it uses may be slower than the one Grok uses — some kind of privileged one, I’m theorizing. Two: the harness instructions and tools make the model use more tokens. It’s probably a mix of both reasons.

The benchmark

What are these results on, anyway? Well, Superconductor enables you to build your own “Custom SWE-Bench”, based on gold standard pull requests into your own codebase. Each pull request is analyzed, the spec is given to the agents you’re benchmarking, and then multiple AI judges grade each agent’s implementation against the reference implementation from the pull request.

How Superconductor Custom SWE-Bench works: select exemplary PRs, select agents, compare the results

The results I’m showing you are on the Custom SWE-Bench we built on our Ruby on Rails codebase, using ten different pull requests. So we stand by Grok being great at Rails code, but if you’re writing COBOL or QBasic, your mileage may vary — you’ll have to sign up, import some PRs, and find out!

If you sign up here and talk to us, we’ll cover the cost of a few agents 🙂.

Closing thoughts

After seeing these results, I’ve personally switched my daily driver to Grok 4.6 via the Grok harness. I was running Grok 4.5 for a while, then recently switched back to GPT 5.6 Sol, but now I’m switching back.

I still like and use Opus and Fable when I need a little extra creativity or persistence. For example, when doing research, planning, or some hard debugging task, I’ll launch several agents, and one of them will always be an Opus or Fable.

But for most dev tasks, I just want a single agent that’s smart, reliable, and fast, and right now, for me, that’s Grok.