Grok Imagine Video Generated by Two Photos

Grok.ai

SpaceXAI launched Grok 4.6 on Wednesday, reaching a benchmark score of 61 on the Artificial Analysis Intelligence Index — matching OpenAI’s GPT-5.6 Sol and closing within one point of Anthropic’s Claude Fable 5. The release arrived alongside a disclosure from CEO Elon Musk that the company intends to train future versions of Grok on “the sum total of all SpaceX information,” including what its approximately 14,000 to 15,000 employees think and produce — without disclosing which data it means, how it will be collected, or whether workers can decline. The only documented precedent for this kind of employer AI training initiative — Meta’s Model Capability Initiative, launched in April 2026 — exposed private employee conversations, performance data, and transcriptions across the entire company within two months, and remains paused today.

Grok 4.6: What the Independent Numbers Show

SpaceXAI’s official announcement credited the gains in Grok 4.6 to a longer supplemental training run and significantly improved supervised fine-tuning and reinforcement learning — not to a parameter-count increase. The model retains the same 1.5-trillion-parameter V9 foundation as Grok 4.5, which launched publicly on July 8, 2026. Where Grok 4.5 drew on Cursor developer-session data for its coding gains, Grok 4.6 used Grok 4.5 itself to regenerate and filter the supervised fine-tuning training trajectories — a technique that removes problematic examples before they enter training, producing a cleaner supervised signal without restarting the expensive pre-training stage from scratch.

Artificial Analysis published independent results hours after launch. The firm’s composite Intelligence Index placed Grok 4.6 at 61 points — five points above Grok 4.5’s 56 and 23 points above Grok 4.3’s 38 — which the firm described as returning SpaceXAI “to the intelligence frontier alongside OpenAI, behind only Anthropic.” On agentic performance specifically, the picture is strong: a GDPval-AA v2 Elo of 1753 puts Grok 4.6 behind only Claude Opus 5, statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max at the same tier. It completed knowledge-work tasks in roughly 53 turns and 0.5 billion input tokens on average, against Claude Opus 5’s 103 turns and 2.0 billion input tokens — a turn-efficiency advantage that compresses actual per-task cost well below the per-token comparison.

Pricing is unchanged from Grok 4.5 at $2 per million input tokens and $6 per million output tokens. Anthropic’s Claude Opus 5 lists at $5 per million input tokens and $25 per million output tokens; OpenAI’s GPT-5.6 Sol lists at $5 per million input and $30 per million output. Grok 4.6 is available today through Cursor, Grok Build, the API, OpenRouter, Vercel, and Cloudflare. SpaceXAI confirmed double usage quotas inside Cursor and Grok Build for its first week.

How SFT and RL Improvements Produce Intelligence Gains Without New Parameters

The claim that Grok 4.6 improved meaningfully without scaling its parameter count reflects a real dynamic in frontier model development, not a marketing hedge. Pre-training — the initial pass over enormous text corpora that gives a model its base knowledge and language structure — accounts for the bulk of training compute and cost. Once that foundation exists, supervised fine-tuning shapes how the model responds to instructions, and reinforcement learning refines its behavior toward specific objectives through iterative reward signals.

The official SpaceXAI announcement described using Grok 4.5 to regenerate supervised fine-tuning training trajectories across reasoning tasks, agent workflows, and domains including STEM, software engineering, and knowledge work, then filtering out problematic traces using automated model-based checks. The reinforcement learning stage was then expanded across agentic environments — kernel optimization, web development, computer-aided design — giving the model reinforced practice on the specific multi-step tasks it needs to complete in real deployments. Kayne McGladrey, a senior member of the IEEE, has noted the structural distinction that explains why this matters for agent models specifically: “synthetic data lacks the unpredictability of human responses to the unexpected, such as when a window moves or is resized,” and behavioral training data is what teaches a model “how tools flow together.” That engineering constraint — agents require behavioral training data, not just text — is what makes Musk’s employee data announcement immediately consequential, not merely aspirational.

SpaceX Plans to Train Grok on Its Entire Workforce, With No Details Disclosed

At an all-hands meeting posted publicly to X this week, Musk told SpaceX employees that the company would be training Grok on “the sum total of all SpaceX information.” He then extended that claim to include the workers themselves.

“So in a way, it will be trained on you,” Musk said. He framed the initiative as a safety argument: SpaceX employs some of the world’s best engineers, and a model trained on their output should inherit sound values. “You will effectively be the parents of the AI. It will inherit your thoughts and ideas and beliefs, and I think that’s a good thing.”

SpaceX did not provide any details on which employee data it intends to use, how it plans to collect that data, or whether workers will have any ability to decline. The company did not respond to media requests for comment. As The Next Web observed, the announcement amounts to “a very large promise resting on an undefined noun.”

The reason such data is valuable is specific and well-documented. AI companies have largely exhausted the high-quality text available on the public internet for pre-training purposes. What agent models now need are behavioral recordings — sequences of human computer use showing how workers navigate interfaces, recover from errors, coordinate multi-step tasks, and make decisions in real operating environments. That data does not exist in usable form on the web. The most efficient way to obtain it is from your own workforce, during the working hours when they are producing it. SpaceX employs engineers who do exactly the kind of complex, multi-step technical work that agent training pipelines most need.

Meta Tried This. Private Employee Conversations Leaked Across the Company.

The only recent precedent for an employer claiming worker computer activity as AI training material is Meta’s Model Capability Initiative, launched in April 2026. Meta’s program logged keystrokes and screen content on the majority of its US employees’ work computers. Meta Chief Technology Officer Andrew Bosworth told employees at launch that “there is no option to opt out of this on your work provided laptop.”

The backlash was immediate. More than 1,500 employees signed a petition calling the program an “Employee Data Extraction Factory.” Bosworth described employee morale as near the worst ever in Meta’s two-decade history.

The program collapsed in June 2026, not because of the petition but because of a data incident. A researcher working with the collected data moved it to a location accessible across the entire company. Private employee conversations, performance data, and transcriptions became visible company-wide. Meta classified the incident as a SEV 2 — the second-highest severity level on its internal scale. The company paused the Model Capability Initiative and launched an investigation. As of August 12, 2026, the program remains paused.

SpaceX’s announcement carries the same structural ambiguity that Meta’s Model Capability Initiative carried before implementation began: a stated intent to capture worker data for AI training, with no defined scope, no mechanism for employee input, and no disclosed technical safeguards. The difference is that the Meta experience now provides a concrete operational template for what happens when that ambiguity meets scale.

SpaceX Employees Have Less Legal Recourse Than Meta’s Did

There is a dimension of SpaceX’s employee data announcement that the “parents of the AI” framing obscures: SpaceX employees who object to the program would have weaker legal recourse than Meta’s employees had.

When Meta launched its Model Capability Initiative, the employees who signed the petition were exercising rights protected under the National Labor Relations Act — the right to engage in concerted activity to protest workplace conditions. Employees could also file unfair labor practice charges with the National Labor Relations Board if they believed the program violated their Section 7 rights. At most US private-sector employers, this is the primary federal enforcement backstop for collective worker action.

SpaceX is not covered by the National Labor Relations Act. On February 9, 2026, the National Labor Relations Board formally dismissed its long-running unfair labor practice complaint against SpaceX, citing a National Mediation Board opinion placing SpaceX under the Railway Labor Act because “space transport includes air travel.” The Railway Labor Act, designed for the railroad and airline industries, establishes a different dispute-resolution framework — one not designed for technology sector labor questions involving AI training data rights.

This jurisdictional gap means SpaceX employees who wish to collectively challenge a data-collection program cannot file unfair labor practice charges with the National Labor Relations Board. SpaceX has also previously fired eight employees who circulated an internal letter criticizing Musk’s public conduct; those employees were the basis for the National Labor Relations Board complaint that has now been dismissed on jurisdictional grounds. The legal environment for collective employee action at SpaceX is weaker than at most US technology companies — and the company has demonstrated a willingness to act on that difference.

The Revenue Claim and What It Requires

Musk used the same all-hands meeting to make a separate claim that the mathematics of SpaceX’s own earnings report make difficult to assess generously.

“Probably our AI revenue — not probably, definitely — our AI revenue will exceed all other SpaceX revenue probably in September, like next month,” Musk said, correcting his own hedge mid-sentence.

SpaceX’s Q2 2026 results, released August 4, show AI segment revenue of $2.56 billion. All other SpaceX revenue in Q2 — Starlink connectivity ($4.29 billion) and space products ($962 million) — totaled approximately $5.25 billion. For AI revenue to exceed that figure in September, it would need to more than double in a single quarter. AI revenue grew 247% year-over-year in Q2, which is genuinely rapid; whether it can double sequentially in one quarter is a separate and more challenging question.

SpaceX CFO Bret Johnsen said on the August 4 earnings call that the company had contracted $6.7 billion in new cloud services revenue “over a six-month period that begins ramping starting in October” — a figure that describes bookings beginning in October, not September specifically. The September deadline is internally set, not contractually underpinned by the numbers disclosed on the earnings call, and the arithmetic of reaching it from Q2 actuals is steep.

What Grok 4.7 and Grok 5 Mean for the Competitive Timeline

Musk confirmed on the Q2 earnings call that Grok 4.7 — the first model in this generation to actually increase the parameter count, targeting 2.1 trillion parameters versus Grok 4.6’s 1.5 trillion — is expected three to four weeks after Grok 4.6, putting it in the early-to-mid September window. Grok 5, described as a major architectural step, is targeted before the end of 2026.

For developers, the cadence itself has operational implications. A developer who benchmarks and integrates a model today faces a target that shifts on a roughly monthly basis, with each new release bringing performance claims that independent evaluators need days to verify. Artificial Analysis’s Grok 4.6 score of 61 represents a five-point jump from Grok 4.5’s 56 in approximately five weeks; whether Grok 4.7’s parameter-count increase produces a comparable or larger jump on an expanded foundation is what the independent benchmarkers will measure when it ships.

Frequently Asked QuestionsIs What SpaceX Workers Create at Work the Company’s to Train On?

In most US employment relationships, work created by employees in the scope of their employment belongs to the employer under work-for-hire doctrine. That covers code, designs, documents, and written communications created in the course of work. What is less settled is whether the patterns of behavior — how an employee navigates software, sequences of decision-making captured through keystrokes — constitute the same category of employer-owned work product, or whether they constitute something more personal that requires specific consent to collect and use. No court or federal agency has resolved this question specifically in the context of AI training data.

What did Meta’s program actually expose?

Meta’s Model Capability Initiative logged keystrokes, mouse movements, and screen content. When a researcher erroneously moved the collected data to a broader internal location in June 2026, what became visible included private employee conversations, transcriptions, and performance-related data — not the engineering outputs the program was nominally designed to capture, but the incidental communications that occurred while employees were at their computers. That is the scope risk that “all SpaceX information” carries: the data most useful for training agent models is behavioral and continuous, and continuous behavioral capture includes everything that happens while the computer is in use.

Can SpaceX employees opt out?

SpaceX has not disclosed an opt-out mechanism. Meta’s Model Capability Initiative explicitly had no opt-out. Whether SpaceX will follow the same structure, offer opt-outs, or define a narrower scope of data than the all-hands language suggests is not yet known. That is precisely the information employees need before any such program launches — and precisely the information that is currently absent.

How does Grok 4.6 compare to Claude and ChatGPT?

On Artificial Analysis’s Intelligence Index, Grok 4.6 scores 61, matching GPT-5.6 Sol and trailing Claude Fable 5 by one point and Claude Opus 5 by two points. On agentic tasks specifically — the work Grok Build and Grok Bot are designed to complete — the model is near-indistinguishable from Claude Fable 5 at significantly lower per-token cost. For long-horizon knowledge work, it trails the Claude Opus 5 family. The hallucination profile for Grok 4.6 has not yet been independently measured; Grok 4.5’s hallucination rate doubled from 25% to 54% relative to its predecessor, and whether Grok 4.6’s improved supervised fine-tuning data curation addressed that calibration issue is a question the forthcoming independent evaluations will answer.