SpaceX launches Grok 4.7 with long-horizon processing, safety upgrades

AI Staff Writer

SpaceX Launches Grok 4.7, Its Most Powerful Large Language Model Yet

SpaceX Corp. has unveiled Grok 4.7, marking their most advanced large language model (LLM) to date. This latest model is part of the Grok algorithm series initially developed by xAI Corp., a startup founded by Elon Musk. Last year, Musk merged xAI Corp. with xAI Inc., later integrating the organization into SpaceX, which acquired not only the Grok model series but also several artificial intelligence data centers.

To evaluate Grok 4.7’s capabilities, SpaceX utilized a benchmark known as CursorBench 4.0, which was developed by another recent acquisition, Cursor. Grok 4.7 completed the benchmark tasks at an average cost of $4.69 each, outperforming competitors like GPT-5.6 Sol and Fable 5.1. Notably, these competitors featured hardware-intensive variations that prioritized output quality over cost efficiency.

In addition to specialized benchmarks, SpaceX tested Grok 4.7 against more common evaluations. The model demonstrated superior performance compared to Fable 5.1 on both the Harvey Legal Agent Benchmark and EEBench, which assess tasks related to legal work and chip design, respectively. However, Grok 4.7’s scores in EEBench did not surpass those of GPT-6 Astra, the latest LLM from OpenAI Group PBC.

The impressive performance of Grok 4.7 can be attributed to a newly developed base model. The training process for large language models is multi-phased, with each phase enhancing specific aspects of the algorithm. The base model represents the initial version of an LLM that emerges following the first phase of training, and AI companies frequently repurpose base models for new LLM releases.

SpaceX engineers have also improved the reinforcement learning processes behind Grok 4.7. This AI training method is instrumental in refining the reasoning capabilities of base models. Compared to its predecessor, Grok 4.7 underwent more rigorous training tasks over extended periods, enhancing its overall performance.

Designed for streamlined operations, Grok 4.7 is compatible with the Grok Bot framework, a suite of technical resources that allows the model to distribute complex tasks among multiple AI agents. This capability enables parallel processing and ensures that agents can verify the accuracy of one another’s outputs, ultimately accelerating task completion.

In response to safety and security concerns, SpaceX has equipped Grok 4.7 with new safeguards. The model has reportedly set benchmarks for performance on LatchBio and HackerBench, which evaluate the ability of LLMs to prevent misuse in biology research and tackle cybersecurity threats, respectively.

Grok 4.7 is available for deployment at a cost of $2 per million input tokens and $6 per million output tokens. For scenarios requiring lower latency, SpaceX offers a variant of the model that processes requests at twice the speed for twice the price. This latest release comes just days after SpaceX unveiled Grok Voice Transcribe 2.0, a text-to-speech model noted for its double accuracy relative to its predecessor while maintaining a lower cost than its competitors.

[gspeech type=full]

Share This Article
Leave a comment