Capacity became visible at every scale. Anthropic tied 220,000 GPUs directly to customer limits; a midtraining experiment encoded behavioral principles before examples; Gemma drafters accelerated local decoding; TSMC contracted 30 years of power; and DeepMind chose an offline virtual economy to test long-horizon agents without exposing live players.
1. Anthropic takes all capacity at SpaceX's Colossus 1 site
Anthropic said its SpaceX agreement would provide more than 300 megawatts of capacity and over 220,000 Nvidia GPUs within a month. It simultaneously doubled Claude Code's five-hour limits, removed peak-hour reductions for Pro and Max users, and raised Claude Opus API limits.
The capacity figures and service effects are Anthropic's claims, but the linkage is unusually explicit: infrastructure supply translated directly into customer quotas. The deal also deepens provider concentration because one competitor's data-center asset now supports another lab's models, despite their rivalry elsewhere.
Sources: Higher usage limits for Claude and a compute deal with SpaceX · Anthropic raises Claude Code usage limits
2. Model Spec Midtraining targets alignment generalization
Anthropic researchers introduced Model Spec Midtraining, an added stage between pretraining and alignment fine-tuning. Models first learn from synthetic documents explaining a behavioral specification, then receive examples of the desired behavior, with the aim of teaching both principles and their application.
On experimental Qwen models, the authors reported agentic-misalignment reductions from 68% to 5% and from 54% to 7%. These results come from synthetic training and controlled evaluations, not deployed Claude behavior, so robustness under different models and real agent environments remains open.
Sources: Model Spec Midtraining
3. Gemma 4 drafters trade extra machinery for faster local inference
Google released experimental Multi-Token Prediction drafters for Gemma 4. The smaller models propose several future tokens, which the main model verifies in parallel; Google reported speedups of 2.8 and 3.1 times on selected Pixel configurations and 2.5 times on Apple M4 hardware.
Because Gemma verifies accepted tokens, speculative decoding is lossless in output distribution, but acceptance and speed vary with hardware and workload. The relevant profile includes prompt length, quantization, memory bandwidth, draft acceptance, and serving framework; Google's peak Pixel and M4 figures describe selected configurations rather than a portable multiplier.
Sources: Gemma 4 models use speculative decoding for faster inference
4. TSMC commits to a 30-year offshore wind purchase
TSMC signed a 30-year agreement for all power from Taiwan's Hai Long offshore wind project, covering more than one gigawatt across three sites. The project is expected to become fully operational in 2027, amid rising chip demand and pressure on Taiwan's imported-energy supply.
Semiconductor capacity planning is now inseparable from power availability. TSMC consumed nearly 10% of Taiwan's electricity in 2023, and cited estimates put that share near one quarter by 2030, making long-term energy contracts a supply-chain input rather than a peripheral sustainability measure.
Sources: TSMC taps wind power as AI chip demand rises
5. DeepMind chooses an offline EVE economy for long-horizon agent tests
Google DeepMind took a minority stake in EVE Online developer Fenris Creations and described the game as a research environment for planning, memory, and continual learning. Experiments will run on a specially designed offline server rather than the live player universe.
The separation is the meaningful safety choice. EVE supplies markets, coordination, long horizons, and strategic adaptation without granting an agent access to a production company or active player economy. Results will describe the custom server's rules and population; transfer to open human environments remains a separate experiment.
Sources: Google DeepMind partners with EVE Online for AI model testing