The week of July 29, 2026, was dominated by a single event: SpaceX’s first-ever public earnings call on August 4, which pulled back the curtain on the financial mechanics of the SpaceXAI era. Elon Musk used the call to outline an aggressive Grok roadmap — Grok 4.6 next week, Grok 4.7 in three to four weeks, and Grok 5 before year-end trained on the entire corpus of SpaceX’s historical data. Away from the earnings spotlight, Grok Voice Think Fast 2.0 completed its production migration, and Grok Imagine Video 1.5 shipped a major feature update adding text-to-video generation and native 1080p output.

Here is our professional analysis of the developments that mattered this week.


SpaceX’s First Earnings Call: $7.8B Revenue, $2.56B AI, and $18.4B Capex

SpaceX held its inaugural quarterly earnings call on August 4, 2026 — the first public financial disclosure since its June 12 IPO under the ticker SPCX. The numbers were striking:

  • Q2 2026 revenue: $7.8 billion, up 92% year-over-year, beating analyst consensus of ~$6.8–6.9B
  • AI segment revenue: $2.56 billion, up 247% YoY, including xAI and Grok subscriptions
  • Adjusted EBITDA: $3.5 billion, up 191% YoY
  • Net loss: $541 million, narrowed from ~$1B a year earlier
  • Capital expenditures: $18.4 billion in a single quarter, up nearly sevenfold from $2.8B a year ago — driven overwhelmingly by AI infrastructure

The market’s reaction was negative. Shares dropped approximately 7% in after-hours trading, largely on the capex figure. Two days later, on August 6, a lockup expiration releases roughly 1 billion insider shares — approximately $106.4 billion worth at Monday’s closing price — with additional tranches unlocking every two weeks until 100% are unlocked by December 2026.

AI as a Formal Business Segment

SpaceX now reports three segments: Space (launch), Connectivity (Starlink), and AI. The AI division — which encompasses xAI, Grok subscriptions, and compute leasing — generated $2.56B in a single quarter. For context, that annualizes to roughly $10B+ in AI revenue, putting SpaceXAI’s AI business in the same revenue neighborhood as established enterprise AI vendors.

Gwynne Shotwell, SpaceX’s President and COO, revealed on the call that token consumption tripled after the release of Grok 4.5 in July. That metric — raw model usage — is the strongest indicator xAI has shared to date that Grok adoption is accelerating materially, even as the user growth plateau at 117 million monthly active users (reported in late July) suggests consumer acquisition has stalled.

The Capex Story: Nvidia Exclusive, 2 GW by Year-End

Musk confirmed on the call that SpaceX will build its AI data centers exclusively on Nvidia’s Vera Rubin architecture, calling it “currently the best AI computer on the market.” The compute roadmap is aggressive:

  • Current online compute: ~1.4 GW
  • Target by end of 2026: 2 GW
  • Target by end of 2027: 10–15 GW

Musk stated that the bottleneck has shifted from GPU supply to power and cooling infrastructure, and that resource investment will focus there going forward.

The Starmind satellite program adds a space-based dimension: SpaceX has co-developed AI computing payloads with Nvidia for its Starmind AI 1 satellites, with each satellite carrying Nvidia Rubin GPUs and Vera CPUs. The first batch is scheduled to begin launching next year, extending AI compute into low Earth orbit.

Grok-Powered Investor Q&A Platform

In a move that drew significant attention, SpaceX bypassed the industry-standard SAY platform (used by Tesla for its earnings calls) in favor of a bespoke investor Q&A tool built with Grok. The platform accepted shareholder questions ahead of the call, with Grok handling deduplication, topic clustering, and ranking.

This is arguably the most visible enterprise deployment of Grok to date — putting the model in front of institutional analysts, retail shareholders, and financial media simultaneously. If the tool performs cleanly, it creates a case study that any other public company could adopt. Tesla’s own earnings calls, which currently rely on SAY, become an obvious next candidate.

(Axios, Business Insider, TradingKey, Basenor)


Grok 4.6 Targeted for Next Week — Grok 5 to Incorporate All SpaceX Data

The earnings call served as a platform for Musk to lay out the most detailed Grok roadmap to date:

  • Grok 4.6: Expected “probably next week” (week of August 7)
  • Grok 4.7: “About three or four weeks from today”
  • Grok 5: “Before the end of the year”

Grok 4.6 is described as a 1.5-trillion-parameter model with significantly improved supervised fine-tuning (SFT) and reinforcement learning (RL) over Grok 4.5. Secondary reporting has produced conflicting parameter counts — some sources reference a “2T model” — but the most consistent characterization from Musk’s own statements is 1.5T with training improvements, not a larger architecture.

Grok 4.7 is reported as a 2.1-trillion-parameter model that will be “better in every way except slightly slower to serve,” with higher token efficiency. This aligns with the previous week’s reporting.

The headline item from the call was Musk’s statement on Grok 5’s training data:

“We will be incorporating the entire corpus of SpaceX data… all the data SpaceX has ever produced over the course of a quarter century.”

This means rocket launch logs, telemetry, manufacturing data, satellite operations data, and proprietary engineering datasets will be fed into Grok 5’s training pipeline. This is the clearest strategic signal yet that the SpaceXAI consolidation is about data synergy, not just corporate restructuring — SpaceX’s proprietary engineering data becomes a durable competitive moat for Grok’s AI capabilities.

The API Reality Check

As of August 5, Grok 4.5 remains the latest documented model in SpaceXAI’s official docs. There is:

  • No public model ID for Grok 4.6 on docs.x.ai
  • No pricing row in the API documentation
  • No context window specification
  • No benchmark scores

The real signal that Grok 4.6 has shipped will be a grok-4.6 entry in the model listing at docs.x.ai/developers/models. LMArena announced on July 29 that Grok 4.6 will appear on the Arena leaderboards the week following its launch, which would provide independent benchmark validation.

(Investing.com transcript, AIBase, Orcarouter)


Grok Voice Think Fast 2.0 Completes Production Migration

On August 5, 2026, the grok-voice-latest alias automatically switched from Grok Voice Think Fast 1.0 to Grok Voice Think Fast 2.0, completing a migration plan xAI announced on July 29. Any developer not explicitly pinned to grok-voice-think-fast-1.0 is now running 2.0.

Specifications

MetricThink Fast 1.0Think Fast 2.0
AA STS Quality Index75.7%82.9%
Agentic τ-voice Bench52.1%56.5%
Time to first audio1.25s0.70s
Reasoning tokens (P50)Baseline~40% of 1.0
Price per minute$0.05$0.08

The model is an end-to-end speech-to-speech system — audio in, audio out — with no separate ASR → LLM → TTS pipeline. It supports 24+ languages, handles noisy telephony conditions, and reasons while speaking with no added delay. Tool calls can begin before the first sentence ends.

Transcription accuracy improved 1.4× over 1.0 and 1.5–2× over leading dedicated STT models (Deepgram Nova 3, ElevenLabs Scribe v2) on xAI’s 24-language benchmark. In noisy telephony conditions, the improvement reaches up to 10×.

Competitive Positioning

At $0.08 per minute ($4.80 per hour), Think Fast 2.0 is approximately 45% of the price of OpenAI’s GPT-Realtime-2.1 High. On the Artificial Analysis Speech-to-Speech Index, Grok Voice TF 2.0 scores 82.9% versus GPT-Realtime-2.1’s 79.1% and Gemini 3.1 Flash’s 69.5%. On the agentic τ-voice benchmark, Grok leads at 56.5% versus OpenAI’s 45.7%.

For enterprise teams building voice agents — call centers, IVR systems, conversational AI — Think Fast 2.0 is now the benchmark leader on paper. The price increase from $0.05 to $0.08 per minute (60% higher) is meaningful for high-volume deployments, but the 60% reduction in reasoning tokens partially offsets the per-minute cost by reducing the compute load per interaction.

Migration Notes

Developers should:

  1. Pin grok-voice-think-fast-1.0 in production if not ready for 2.0
  2. Re-cost at $0.08/min versus prior $0.05 assumptions
  3. Enable binary audio transport if controlling both endpoints
  4. Add keyterms for product names and replace for pronunciation corrections
  5. Implement tool → wait-for-playback → response.create ordering to avoid audio overlap
  6. Use ephemeral tokens for any client that is not your server

The WebSocket endpoint is wss://api.x.ai/v1/realtime?model=grok-voice-think-fast-2.0. Codecs supported: PCM (8–48 kHz), μ-law, A-law, and Opus (24 kHz).

(xAI docs, Explainx, Artificial Analysis)


Grok Imagine Video 1.5: Text-to-Video, Native 1080p, Voice Reference

On July 31, 2026, xAI shipped a significant update to Grok Imagine Video 1.5, adding four capabilities the platform had been missing since its May launch:

Text-to-Video Generation

Users can now generate videos from text prompts alone without supplying a reference image. The implementation pairs xAI’s image generation pipeline with the image-to-video engine — the model synthesizes a starting frame from the text prompt, then animates forward. This closes a feature gap with Google Veo 3.1 and Seedance 2.0, both of which have offered text-to-video at the API level since their launches.

Native 1080p Output

Previous Grok Imagine Video output was capped at 720p (with a brief “1080p toggle” that was actually upscaled 720p). The update delivers true native 1080p rendering. The pricing reveals the architectural cost: Aurora is an autoregressive system that generates each frame sequentially, so scaling from 720p to 1080p multiplies pixel tokens by ~2.25×. API pricing reflects this:

ResolutionPrice per second
480p$0.08
720p$0.14
1080p$0.25

At 1080p, generating 60 minutes of video via the API costs $900 — compared to $504 at 720p and roughly $540 at Google Veo 3.1 Fast tier pricing. The 1080p capability closes the format gap with competitors; it does not close the cost gap at that resolution.

Voice Reference: Face and Voice Consistency

The most distinctive new feature: users supply a character photo and a voice sample, and the model maintains both face and voice across every generated scene. The underlying mechanism is a speaker embedding system that extracts voice identity characteristics (tone, pitch, prosody, accent) separately from linguistic content.

This removes a post-production step that traditionally required combining separate tools. HeyGen’s Avatar API, for comparison, charges $3 per minute for talking-head clips. Grok Imagine’s voice reference is embedded in full-scene video generation — not just avatars — representing a different capability category.

Access: Voice reference is rolling out first to SuperGrok Heavy and SuperGrok Plus subscribers in the United States, with broader rollout indicated within days. API developers must contact xAI’s sales team to request access.

Consent gap: xAI’s announcement does not include any discussion of consent guardrails. The feature requires a photo and voice sample but does not appear to require that they belong to the same person, or that the person depicted has consented. Given Grok’s history with non-consensual imagery — an estimated 3 million sexualized images generated in an 11-day window in late 2025/early 2026, followed by federal lawsuits and investigations from 35 state attorneys general — combining face and voice reference in a single generation call creates a more complete non-consensual intimate video capability than image generation alone. xAI’s acceptable use policy prohibits such content, but enforcement history gives that prohibition limited reassurance without disclosed technical controls.

Multi-Reference Scene Control

Up to seven simultaneous reference images can be used in a single generation. Each reference locks one visual element — a face, a product, a location, a prop — while allowing other elements to vary. This is a larger anchor count than most dedicated reference tools offer and, paired with voice reference, produces a character pipeline where face, voice, and environmental context can all be locked simultaneously.

API availability: Text-to-video, native 1080p, and image references are available now with a standard API key. The grok-imagine-video-1.5 model string replaces the earlier -preview alias. Voice reference requires sales team access.

(TechTimes, xAI news, Scenario help)


Competitive Landscape: OpenAI GPT-5.6 Sol, Terra, and Luna

OpenAI shipped GPT-5.6 on July 9 as a three-model family — the most significant competitive release during this period:

ModelPositioningPrice (per 1M tokens)
SolFlagship, highest capability$5 input / $30 output
TerraBalanced, GPT-5.5-level at lower cost$2.50 input / $15 output
LunaFastest, cheapest, high-volume$1 input / $6 output

OpenAI also introduced Ultra mode — its highest-capability setting that coordinates multiple agents across parallel workstreams — and a Max reasoning level available in ChatGPT Work and Codex.

The three-tier strategy contrasts sharply with xAI’s approach. SpaceXAI has been consolidating its API around a single flagship model (Grok 4.3, then 4.5) and using subscription tiers for distribution rather than model tiers. OpenAI’s Sol/Terra/Luna split gives buyers a price-performance tradeoff within a single model family — a procurement advantage for organizations that need different capability tiers for different workloads.

Meanwhile, Moonshot AI’s Kimi K3 — a 2.8-trillion-parameter open-weight model with 896 experts in a Mixture-of-Experts architecture and a 1-million-token context window — was released in late July, claiming the title of largest open-weight model ever. The competitive pressure on xAI is intensifying from both proprietary and open-source directions.


What to Watch

  • Grok 4.6 launch confirmation: Watch for the model ID appearing in xAI’s API docs and LMArena leaderboard. The gap between “next week” and actual API availability will indicate how much lead time to build into planning. Musk’s delivery cadence is accelerating but his dates are targets, not commitments.
  • SpaceX lockup cascade: August 6 begins the first tranche of insider share unlocks — roughly 1 billion shares, ~$106B. Additional tranches every two weeks through December. The selling pressure will shape SPCX’s stock performance and, by extension, the market’s appetite for AI infrastructure spending at this scale.
  • Grok 5 and SpaceX data integration: The promise to train on “all data SpaceX has ever produced” is strategically significant. If SpaceX’s engineering telemetry, manufacturing data, and operational logs produce measurable capability gains in Grok 5, it validates the entire SpaceXAI consolidation thesis. Watch for any technical disclosures about data pipeline architecture.
  • Voice Think Fast 2.0 production performance: The benchmark numbers are strong on paper. Enterprise deployments will reveal whether the 60% token reduction and 0.70s time-to-first-audio translate to real-world cost savings and user experience improvements. Watch for third-party evaluations and case studies.
  • Grok Imagine consent guardrails: Voice reference combined with face reference creates non-consensual intimate video risk. Watch for whether xAI discloses technical controls — not just policy prohibitions — before broader rollout.
  • Starmind satellite launches: The first batch of AI-equipped satellites is scheduled for next year. If successful, space-based AI compute becomes a differentiator no competitor can replicate.
  • OpenAI competitive response: GPT-5.6 Sol scores 80 on the Artificial Analysis Coding Agent Index. Grok 4.6’s benchmark numbers — when they arrive — will be directly compared. The three-tier Sol/Terra/Luna pricing model may force xAI to reconsider its single-flagship API strategy.

Follow Kevin Kaminski on X for daily AI ecosystem updates. Check back next week for the next xAI Weekly.