Theo is a developer and founder heavily involved with agentic coding tools, T3 Code, Cursor integrations, and AI model testing. He reviews xAI’s newly released Grok 4.6. He notes xAI has accelerated dramatically since acquiring Cursor, moving from long dry spells to rapid back-to-back releases.
Grok 4.6
It is not a new pre-train but a substantial post-training / RL upgrade on Grok 4.5, leveraging Cursor’s post-training and RL expertise.
Strong focus on long-running agents, multi-step reliability, sub-agents, self-verification, and staying on complex tasks like research, codebases, turning ideas into working apps.
Benchmarks are essentially neck-and-neck with GPT-5.6 Sol and just behind Fable 5 and Opus 5.
Meaningful gains over Grok 4.5 on DeepSuite, Cursor Bench, Frontier Code, AA Briefcase, Harvey Lab.
XAI Grok 4.6 is now firmly in the frontier pack on many agentic benchmarks.Theo is not yet ready to daily-drive it over Fable/Sol for most work, and design/game generation remain weak spots. He is cautiously optimistic about the agentic direction and the rapid pace, mixed on the practical regressions, and clear that real-world coding usefulness is higher than pure design or creative 3D tasks.
It feels fast, relatively cheap, and reliable for agentic work. The official Grok Build CLI is excellent.
Available immediately in Cursor and Grok Build (with temporary 2× usage).
Strong agentic scores at lower cost than the top closed models.
Elon has already teased that Grok 4.7 is coming in 3–4 weeks and is significantly better.
Pricing & efficiency (notable regression)$2 / million input tokens, $6 / million output.
Real-world cost per task ~$0.84 (competitive with cheaper models, far cheaper than Opus/Fable).
However, it uses >30% more tokens than Grok 4.5, so it is meaningfully more expensive and slower than its predecessor (no longer the ultra-efficient “magic” model). Cache reads also got slightly more expensive. Context window remains 500k.
Real-world testingDesign / UI generation
Disappointing. Outputs feel dated / “old AI slop” (noisy backgrounds, sharp text, Tailwind-templaty cards, brutalist layouts). Noticeably worse than Fable, Opus, or even Soul. Claude’s design skill helps a bit, but still not competitive.Game ports (“Fish Slop” – an aquarium-style game the reviewer originally built with Opus) 2D version: usable but rough (bad controls, wrong sizing, poor pacing, UI bugs).
3D version: first model in a long time to outright fail on the initial attempt (black screen). After a screenshot fix it ran, but results were poor (inverted axes, bad placement, low-quality models). Open-weight models like Kimmy K2 and even Muse produced clearly superior 3D results.
Serious coding / agentic work
Much stronger. Solid security audit of a real project (Lakebed).
Accurate analysis of migrating T3 Code’s Cursor integration from the outdated ACP adapter to the official SDK.
Successfully planned, implemented, filed, and “babysat” a ~1,000-line PR improving Grok Build support in T3 Code, then stacked a second PR based on the reviewer’s local history and event gaps. Handled complex multi-step, multi-context work coherently (one of the areas where Grok 4.5 already impressed him).

Brian Wang is a Futurist Thought Leader and a popular Science blogger with 1 million readers per month. His blog Nextbigfuture.com is ranked #1 Science News Blog. It covers many disruptive technology and trends including Space, Robotics, Artificial Intelligence, Medicine, Anti-aging Biotechnology, and Nanotechnology.
Known for identifying cutting edge technologies, he is currently a Co-Founder of a startup and fundraiser for high potential early-stage companies. He is the Head of Research for Allocations for deep technology investments and an Angel Investor at Space Angels.
A frequent speaker at corporations, he has been a TEDx speaker, a Singularity University speaker and guest at numerous interviews for radio and podcasts. He is open to public speaking and advising engagements.
