Elon Musk's SpaceXAI to Add 660,000 GPUs This Year, Nearing 1.44 Million Total
Elon Musk's SpaceXAI is poised for an unprecedented expansion, adding another 660,000 AI GPUs this year to reach nearly 1.44 million units, a move that solidifies its position at the forefront of the global AI compute race and underscores the escalating demand for massive computational power in the artificial intelligence industry.
✨ This content was summarized and interpreted by AI; it may contain errors — please verify accuracy with the original sources. Learn more
Listen to this story

Elon Musk's SpaceXAI is poised for an unprecedented expansion, adding another 660,000 AI GPUs this year, which will bring its total operational fleet to approximately 1.44 million units and solidify its position at the forefront of the global AI compute race. The core of this expansion centers on the Colossus 2 supercomputing facility, which is slated to receive 220,000 NVIDIA GB300 GPUs by next week, with two additional tranches of the same amount expected to come online later this year. This influx will push Colossus 2 alone to over one million AI GPUs, complementing existing infrastructure that includes 150,000 H100, 50,000 H200, and 30,000 GB200 GPUs at Colossus 1, and an initial 110,000 GB200 and 440,000 GB300 GPUs already at Colossus 2. Powering this colossal computational engine is a monumental undertaking, addressed by the ongoing construction of a 1.2-gigawatt permanent power plant, a necessity highlighted by the controversies surrounding SpaceXAI's earlier reliance on unpermitted gas turbines, which are now slated for removal by July 2027.
This aggressive scaling by SpaceXAI, a company only three years old since xAI's acquisition by SpaceX in February 2026, marks a critical inflection point in the artificial intelligence industry, signifying a relentless pursuit of computational dominance. The deployment of NVIDIA's GB300 GPUs, part of the Blackwell Ultra architecture, is particularly impactful. Each GB300 NVL72 rack integrates 72 Blackwell Ultra GPUs and 36 Grace CPUs, pooling 20 terabytes of HBM3E memory and delivering up to 15,000 dense FP4 TFLOP/s per chip. This advanced architecture, featuring liquid cooling, is engineered to handle the most demanding AI workloads, enabling the training of increasingly complex frontier models and supporting longer contextual windows for applications like Grok. The sheer volume of these high-performance chips will accelerate SpaceXAI's ability to develop, train, and deploy next-generation AI, potentially leading to breakthroughs that are currently compute-limited. It underscores the "compute is king" philosophy driving the current AI arms race, where access to vast computational resources is a primary determinant of innovation and competitive advantage.
However, this unprecedented scale introduces significant challenges, particularly concerning energy consumption. A single GB300 NVL72 rack demands between 132 and 142 kilowatts at nominal load, peaking near 155 kilowatts. Extrapolating this, a million GPUs could collectively demand roughly 1.875 gigawatts of power, a figure comparable to the output of several nuclear power plants. The 1.2-gigawatt power plant under construction is thus not merely an amenity but an existential requirement to bring these systems fully online. Such immense power needs strain existing electrical grids, necessitating substantial infrastructure investments and raising environmental concerns, as evidenced by the legal disputes over SpaceXAI's temporary gas turbines and the requirement for a dedicated clean water recycling plant for cooling. These challenges highlight that the AI arms race is as much an energy race as it is a chip race, pushing the boundaries of data center design and sustainability. The concentration of such massive compute power also raises questions about the future landscape of AI, potentially consolidating control over advanced AI development in the hands of a few well-resourced entities.
Comparing SpaceXAI's ambitions to its rivals reveals the intensity of this global compute arms race. Meta, for instance, aimed to have 1.3 million GPUs in service by the end of 2025, with its AI fleet alone drawing an estimated 1.8 to 2.0 gigawatts of IT load. Earlier, Meta had committed to 350,000 NVIDIA H100 GPUs by the end of 2024. OpenAI, backed by Microsoft, is also a formidable contender, with a total operational and planned GPU count of 1.4 million across nine data centers, boasting a power capacity of 3.7 gigawatts. OpenAI's CEO Sam Altman revealed plans to have "well over 1 million GPUs online" by the end of 2025 and has secured 5 gigawatts of dedicated NVIDIA Vera Rubin capacity. Microsoft itself has deployed hundreds of thousands of liquid-cooled NVIDIA Grace Blackwell GPUs. In contrast, Google has pursued a different strategy, relying heavily on its custom-designed Tensor Processing Units (TPUs), with its latest generations like Ironwood (TPU7x) and Trillium (v6e) offering specialized performance and energy efficiency for AI workloads. Google claims its vertically integrated TPU infrastructure can offer lower capital and operational expenditures for large language model training compared to NVIDIA H100 clusters. SpaceXAI's rapid ascent is particularly striking, given its relatively young age compared to established players like OpenAI (ten years old) and Anthropic (six years old). Furthermore, the strategic decision to rent out capacity from its Colossus 1 site, which was found inefficient for Grok training, to companies like Anthropic for inference, demonstrates an agile approach to resource optimization and commercialization.
Looking ahead, the trajectory of AI development will undoubtedly be shaped by this relentless pursuit of computational scale. Elon Musk's stated ambition to grow SpaceXAI's data center capacity sevenfold by 2027 and target 50 million H100-equivalent GPUs by 2030 signals a future where AI models will demand orders of magnitude more compute than today. However, this growth will be constrained by more than just chip supply. The major bottlenecks will increasingly be power generation, advanced cooling solutions (with liquid cooling becoming standard for high-density racks), and sophisticated networking infrastructure, as even 110,000-chip increments are limited by fiber optic cable capacity. The environmental and regulatory pressures will intensify, pushing companies to invest in sustainable energy sources and advanced water management systems, as seen with SpaceXAI's commitment to its 1.2 GW power plant and water recycling initiatives. Furthermore, the increasing costs and complexity of building and operating such infrastructure might lead to greater consolidation in the AI industry or, conversely, drive more players to explore custom silicon solutions, following Google's lead, to optimize for specific workloads and reduce reliance on a single GPU vendor. The emergence of these hyperscale AI factories, capable of both training and offering AI as a service, will also democratize access to cutting-edge AI for other businesses and researchers, albeit within a highly concentrated and competitive landscape. The "AI arms race" is evolving into a contest of national strategic importance, where compute power, energy independence, and advanced infrastructure are becoming critical metrics of global influence.