The official xAI launch page showing the date August 12, 2026 and the headline Introducing Grok 4.6
xAI presents Grok 4.6 as a model focused on long-running agents and interactive and visual work. Source: xAI; checked 2026-08-24.

Grok Bot arrived on August 11, followed by Grok 4.6 on August 12. Releasing the agent and the model one day apart makes the two announcements more useful together than separately. One supplies a reasoning engine designed for long work; the other supplies a place where that work can continue across a browser and apps.

The eye-catching parts are the benchmark scores and the low API price. Neither proves that a workflow will finish safely. Comparing the official announcements, model documentation, product FAQ, and the relevant SEC filing points to a more consequential change: xAI is shipping the model, execution environment, and operational control layer as one connected stack.

The two launches do different jobs

xAI describes Grok 4.6 as a model focused on long-running agent work plus interactive and visual projects. Grok Bot is an always-on agent that keeps working inside a cloud computer and returns when it needs approval or judgment. Despite its name, it is closer to a remote work assistant than to a conventional support chatbot.

The official xAI page introducing Grok Bot as an agent that works in a cloud computer

Source: xAI, Introducing Grok Bot. Checked 2026-08-24.

xAI says the bot signs in to tools and apps, operates their interfaces, and continues until it has a result. The examples span sales research, marketing preparation, operations, and bug fixes. “Works end to end” should not be read as unlimited autonomy, however. Payments, external messages, deletions, and permission changes remain good places for a human approval gate because they are costly or difficult to reverse.

Access expanded after launch. As of August 21, Grok Bot is included with SuperGrok Plus and Heavy, Cursor Pro+ and Ultra, and Cursor Teams Standard and Premium. The official FAQ lists macOS, Windows, and iOS 18; Linux, Android, and iPad were not supported at the initial launch. Enterprise availability is rolling out separately.

Grok 4.6 is built around longer work trajectories

In xAI’s published table, Grok 4.6 High scores 61 on the Artificial Analysis Intelligence Index, 69.9% on CursorBench v3.2, 65.9% on DeepSWE v1.1, and 61.3% on FrontierCode Extended. Other models lead individual rows in the same table. The defensible reading is that Grok 4.6 is competitive across several long-horizon agent evaluations, not that it wins every benchmark.

xAI’s evaluation table comparing Grok 4.6 High, Grok 4.5 High, GPT-5.6 Sol Max, and Fable 5 Max

Source: xAI, Introducing Grok 4.6. Bold cells indicate the best result in each row; competing-model figures are presented by xAI. Checked 2026-08-24.

These are vendor-supplied results. A buying decision needs a second test built from the organization’s own repositories, browser procedures, document formats, and recovery rules. For a long-running job, the first answer matters less than whether the agent detects an intermediate mistake, retries sensibly, checks its work before finishing, and stays inside a cost ceiling.

Artificial Analysis independently reports the same Intelligence Index score of 61, but its API measurements also put output speed at 65.8 tokens per second and time to the first answer token at 41.87 seconds. That snapshot separates model quality from interactive speed: a strong intelligence score does not guarantee a fast first response. Rankings and latency can change, so these figures are dated August 24, 2026.

The model documentation lists a 500,000-token context window, text and image input, function calling, structured output, and reasoning. Base API pricing is $2 per million input tokens, $0.50 per million cached input tokens, and $6 per million output tokens. A separate high-context tier applies above 200,000 tokens, so support for a 500K context does not mean every token in that range always uses the base rate.

The xAI model documentation showing Grok 4.6 context length, input pricing, cached input pricing, and output pricing

Source: xAI model documentation, Grok 4.6. Checked 2026-08-24.

At launch, Grok 4.6 was available through Cursor, Grok Build, the xAI API, OpenRouter, Vercel, and Cloudflare. Cursor and Grok Build offered double included usage during the first launch week. That was a temporary launch condition, not the permanent unit price.

Grok Bot isolates users, not individual bots

One claim in the supplied material needs a direct correction. Grok Bot does not assign a separate cloud computer to every bot. The official FAQ says that all Grok Bots owned by one user share one persistent cloud computer, including its files, browser, and logins. Isolation is per user, not per bot.

Inbox, research, and operations bots sharing one user-level cloud computer with shared files, browser, and logins, bounded by approval, isolation, and usage controls

That distinction changes the security design. A login created for a research bot and browser state used by an operations bot live inside the same user boundary. Naming two bots differently is not a permission boundary. Sensitive-action approval, user-level isolation, and hard limits on access and usage therefore belong in the operating policy.

xAI documents approval boundaries and says Auto Review applies when enforcement is available. Cursor separately documents encryption, Privacy Mode, and enterprise and network controls. These are available controls, not proof that a particular deployment is safe. The adopter still has to test which actions pass automatically, when a person is called back, whether an external side effect is recorded, and how a failed job is rolled back.

Bot access is priced separately from the model API. The access page now spans several Cursor, SuperGrok, and Teams tiers and has already changed since launch, so a fixed price pair would become misleading. Each eligible plan includes weekly usage and may allow additional token-billed usage. Check the current plan page, then measure total cost per completed task rather than subscription price alone.

The real competition is now the connected work stack

A capable model cannot finish recurring work without somewhere to act. It needs a repository where code can change, a browser where an interface can be operated, an inbox where a message can be read, and a control surface where risky actions can stop. Pairing Grok 4.6 with Grok Bot connects a reasoning engine to that execution layer.

A stack connecting the Grok 4.6 model to Cursor, Grok Build, and Grok Bot, then to code, browser, apps, and inbox, with permissions, approvals, audit, and cost limits

The relationship between Cursor and SpaceX changed on August 14. SpaceX signed the merger agreement with Cursor operator Anysphere on June 16, 2026 at an implied equity value of about $60 billion. A later SEC filing says the merger became effective on August 14 and Cursor survived as a wholly owned SpaceX subsidiary. The earlier filing described a pending transaction; that is no longer the current status.

It would be equally premature to attribute xAI’s position to that agreement alone. What can be observed is a direction: the xAI model, Cursor and Grok Build work surfaces, Grok Bot’s persistent computer, and SpaceX’s computing and capital base are becoming more tightly connected. Whether that becomes an operating advantage must be measured through completion, total cost, permission incidents, and the time people spend repairing agent output.

A four-week work sample is a better adoption test

There is little reason to connect an entire inbox and payment account in week one. Start with 20 to 30 reversible tasks that have unambiguous completion criteria, then score model fit and bot fit separately. Long coding or research jobs test Grok 4.6’s ability to recover and verify. Repeated browser procedures test Grok Bot’s persistence and approval flow.

A four-week pilot scorecard separating model fit, bot fit, and stop signals such as unclear shared-login scope, missing rollback, and missing usage limits

The operating sheet should capture completion rate, human correction time, approval requests, recovery time after failure, and token cost per completed task. It should also record the scope of shared logins and every external send or deletion. If output quality rises while credential or cost control becomes harder to understand, access should not expand.

A practical sequence is read-only research and document drafts in week one, sandboxed browser work in week two, approval-before-action in week three, and a narrow repeated routine only in week four. The success test is not that an agent stayed busy for hours. It is that the work finished safely with less human attention.

Keep the comparison fair by running the same task set through the current process as a control. Record the person’s elapsed time, rework, and failure recovery alongside the agent’s results. Otherwise a polished demonstration can look productive even when it merely moves review work to a later stage.

Grok 4.6 and Grok Bot are evidence that frontier-model competition is becoming work-system competition. Predictions about a new “top three” are less useful than one operational question: do the strong model, the computer where it acts, and the controls that can stop it work together? Only then does an always-on agent become more than another chat window.

References and reporting

Public pages used for reported facts, official documentation, policy background, product details, and claims that may change.