NEWS
Apple Built Intelligence on Google TPUs and Nvidia GPUs
Apple trained Apple Intelligence on Google TPUs, then put its top 2026 model on Nvidia GPUs in Google Cloud.
Apple trained Apple Intelligence on 2,048 Google TPUv5p chips and 8,192 TPUv4 chips, its July 2024 research paper said. By June 2026 the same program still pretrained on Google’s cloud TPUs, and the top model ran on Nvidia GPUs in Google Cloud.
That is a long way from the 2024 reading, which treated the missing Nvidia mention as a clean break. Training stayed on Tensor Processing Units. The hardest serving job did not.
Apple Trained the First Models on Google Cloud TPUs
The July 2024 report, titled Apple Intelligence Foundation Language Models, described two systems. AFM-on-device was a roughly 3-billion-parameter model meant to run on iPhone, iPad, and Mac. AFM-server was a larger model meant for Private Cloud Compute, Apple’s sealed cloud path for jobs that do not fit on the device.
Apple 6.3 trillion tokens on TPUv4 for AFM-server, trained from scratch on 8,192 chips with a sequence length of 4,096 and a batch of 4,096 sequences. Those chips were provisioned as eight slices of 1,024, with only data-parallel work crossing the slower data-center network between slices. AFM-on-device was distilled and pruned from a larger model on one slice of 2,048 TPUv5p chips.
The paper put both runs on v4 and v5p Cloud TPU clusters using AXLearn, Apple’s JAX-based library built for public cloud. It did not describe Nvidia GPUs in that training stack. The split was two different jobs on two TPU generations, not one giant mixed cluster.
THE 2024 AND 2025 TRAINING RUNS
| Run | Chips | Tokens |
|---|---|---|
| 2024 AFM-on-device | 2,048 TPUv5p | Distilled from a larger model |
| 2024 AFM-server | 8,192 TPUv4 | 6.3 trillion |
| 2025 server model | 8,192 TPUv5p | 13.4 trillion |
Apple also wrote that no private Apple user data sat in the mixture. The stack was licensed publisher data, curated public and open-source sets, and pages crawled by Applebot, with robots.txt opt-outs respected.
Google’s TPU Pods Only Run Inside Google Cloud
A TPU is Google’s custom trainer, sold as a cloud service rather than a box a customer racks at home. Google’s own v5p documentation puts 8,960 chips in a v5p pod, each rated at 459 TFLOPS in BF16 with 95 GiB of high-bandwidth memory. A Google Cloud post said v5p can train large language models 2.8 times faster than v4.
That layout is why 2,048 chips reads as a round slice, not a random purchase. Apple’s on-device run filled one 2,048-chip slice. The 2025 server run used four such slices. The chips talk over Google’s inter-chip fabric inside a slice, then over ordinary data-center links between slices, which is the topology Apple’s 2024 paper described.
The software follows the hardware. AXLearn sits on JAX, which is the stack Google’s TPU runtime expects. Renting those chips meant writing the first Apple Intelligence trainers to Google’s cloud, not assembling a private Nvidia cluster running CUDA. The 2024 paper made that path look like a one-time compute buy. It was also an on-ramp into Google’s tools.
8,192 Newer Chips, 13.4 Trillion Tokens
A year later Apple published a 2025 technical report, dated July 17, 2025, for the models shown at WWDC that June. The server model was still a Cloud TPU job. Apple trained it on 8,192 v5p accelerators provisioned as four 2,048-chip slices, again through AXLearn, for 13.4 trillion tokens.
The chip count matched 2024. The generation did not. v5p replaced v4, and the token budget rose from 6.3 trillion to 13.4 trillion. Apple said built-in fault tolerance in AXLearn kept overall good output at 93 percent when hardware hiccuped. The on-device line stayed a roughly 3-billion-parameter model, now squeezed with 2-bit quantization-aware training for Apple silicon.
The 2025 server design, a Parallel-Track Mixture-of-Experts transformer, was written for Private Cloud Compute cost and quality. Training, once more, was a Google cloud TPU problem. If the 2024 paper was a trial rental, the 2025 run was a renewal at larger scale.
Why the Top Model Now Runs on Nvidia GPUs
On June 8, 2026, Apple Machine Learning Research posted the third generation of Apple Foundation Models. At the center, it said, was a family of five foundation models with Google, spanning on-device jobs to Private Cloud Compute. Pre-training, Apple wrote, “significantly scaled” on “the latest generation of cloud TPU accelerators.” The Google training habit held.
Serving split. AFM 3 Core, AFM 3 Core Advanced, AFM 3 Cloud, and ADM 3 Cloud were purpose-built for Apple silicon. AFM 3 Cloud Pro, the most capable server model, was optimized for NVIDIA GPUs. Apple said it worked with Google and NVIDIA to extend Private Cloud Compute to those GPUs in Google Cloud.
THE FIVE AFM 3 MODELS
- AFM 3 Core: Next 3-billion-parameter dense on-device model, built for everyday Apple Intelligence tasks.
- AFM 3 Core Advanced: 20-billion-parameter sparse on-device model that activates 1 to 4 billion parameters per request, stored in flash and paged into DRAM.
- AFM 3 Cloud: Server workhorse on Apple silicon inside Private Cloud Compute.
- ADM 3 Cloud (Image): Image generation and editing model, also on Apple silicon, for Photos tools and Image Playground.
- AFM 3 Cloud Pro: Top server model for agentic tool use and hard reasoning, on NVIDIA GPUs in Google Cloud.
Four of the five still run on Apple’s own chips. The exception is the model Apple itself calls most capable. NVIDIA said its GPUs with Confidential Computing on Blackwell GPUs now handle confidential inference for Apple Intelligence as Private Cloud Compute moves onto Google Cloud.
That is the turn the 2024 coverage missed. Skipping Nvidia for training did not keep Nvidia out of the product. Users never sit on a training cluster. They sit on whatever serves the hard query, and Apple put that slice on Nvidia iron in a rival’s data center.
HUMAN PREFERENCE, 2026 VS 2025
- On-device text: Graders preferred AFM 3 Core on 45.6 percent of prompts, against 23.3 percent for the 2025 baseline.
- Server text: AFM 3 Cloud was preferred on 64.7 percent of prompts, against 8.7 percent for the 2025 AFM Server model.
- Cloud Pro lift: Roughly 10 percent better response satisfaction on text and 14 percent on image understanding than AFM 3 Cloud.
- Voices: AFM 3 Core Advanced scored 4.15 against 3.87 on general speech, and 4.24 against 3.82 on conversational speech, on a 5-point mean opinion scale.
Google’s payoff is easy to overstate as “Siri is Gemini,” which Apple does not say. The published fact is narrower and stickier. Google supplied TPU training in 2024 and 2025, model work in 2026, and the cloud that hosts Cloud Pro. It is the same shape as Search sitting under Safari: the brand on the glass is Apple, the compute underneath is Google’s.
Private Cloud Compute Leaves Apple’s Own Racks
When Apple introduced Private Cloud Compute in 2024, the pitch was that cloud Apple Intelligence still ran on Apple silicon in Apple data centers, with cryptographic checks a phone could verify. In June 2026 the company said it was expanding Private Cloud Compute beyond those rooms for the first time.
Our core PCC requirements remain exactly the same: stateless computation, enforceable guarantees, no privileged runtime access, non-targetability, and verifiable transparency. What’s new with PCC on Google Cloud is the implementation: NVIDIA Confidential Computing with NVIDIA GPUs, Intel CPUs with TDX, and Google’s Titan chip.
Apple Security Research, Expanding Private Cloud Compute, June 2026
Apple said it keeps signing keys, publishes binaries for inspection, and holds a cryptographically verifiable, append-only ledger of Google Cloud hardware in the PCC fleet. Devices, it said, will only trust PCC software Apple has approved. Google Cloud wrote that the serving platform uses Intel TDX and NVIDIA Confidential Computing so the path from CPU to GPU stays inside hardware isolation, with Titan as a hardware root of trust.
Apple also said the Google Cloud deployment would ramp toward the full set of protections during a summer preview. The privacy line on training did not move. The 2026 post repeated that Apple does not use users’ private personal data or user interactions to train foundation models.
The awkward part is physical, not rhetorical. The product that was sold as Apple silicon in Apple rooms now has a top tier that runs on Nvidia GPUs in Google’s rooms, wrapped in the same PCC name. Stateless math can be true on rented chips. It is still rented chips.
A 250,000-Square-Foot Plant for Apple Silicon Servers
Apple is not treating the Google Cloud slice as the whole future of serving. On February 24, 2025, it said it would spend more than $500 billion in the United States over four years, including a 250,000-square-foot Houston plant to assemble servers for Apple Intelligence and Private Cloud Compute. Those servers, previously made outside the United States, were designed around Apple silicon in the data center, with more capacity planned in North Carolina, Iowa, Oregon, Arizona, and Nevada.
The same release doubled Apple’s U.S. Advanced Manufacturing Fund from $5 billion to $10 billion and said the company planned to hire around 20,000 people, most of them in R&D, silicon, software, and machine learning. That $5 billion figure is the manufacturing fund, not a separate Apple Intelligence server budget.
HOW THE HARDWARE CHOICE MOVED
- July 2024: Apple publishes the first AFM paper, with AFM-server on 8,192 TPUv4 chips and AFM-on-device on 2,048 TPUv5p chips.
- February 24, 2025: Apple announces the Houston server plant and a more than $500 billion U.S. spend, with Apple silicon as the data-center pitch for Private Cloud Compute.
- June 9, 2025: Apple shows updated on-device and server foundation models, still aimed at Apple silicon and Private Cloud Compute.
- July 17, 2025: The 2025 tech report puts the new server training run on 8,192 TPUv5p chips for 13.4 trillion tokens.
- June 8, 2026: Third-generation models are built with Google, pretrained on the latest cloud TPUs, and AFM 3 Cloud Pro is served on NVIDIA GPUs in Google Cloud.
Read in order, the sequence is rental, then a plant, then a larger rental, then a hybrid. Houston is the in-house serving bet. Cloud Pro is the admission that Apple silicon in Apple rooms was not enough for the top model in 2026. Training never left Google’s TPU fleet in the papers Apple has posted.
The 2024 TPU disclosure was never a verdict on Nvidia’s business. It was a training invoice. That invoice was paid again in 2025, and in 2026 it sat beside a second invoice for Nvidia inference on Google Cloud, while a Texas plant tried to put Apple silicon back under the rest of the load.
Frequently Asked Questions
What chips did Apple use to train the first Apple Intelligence models?
The 2024 paper used Google Cloud TPUs only in the training section it published: 2,048 TPUv5p chips for AFM-on-device and 8,192 TPUv4 chips for AFM-server. AFM-on-device used 26 transformer layers and a 3,072 model dimension, with about 2.58 billion non-embedding parameters, and Apple said it was distilled and pruned from a larger parent rather than trained from scratch like the server model.
Did Apple use Nvidia GPUs to train those 2024 models?
The 2024 infrastructure write-up names v4 and v5p Cloud TPU clusters and the AXLearn library on JAX, and it gives a sustained model-flop-utilization of about 52 percent for the AFM-server run. Nvidia GPUs show up later, in June 2026, as the serving hardware for AFM 3 Cloud Pro, not as the 2024 trainers.
What is a Google TPU, and how big is a v5p pod?
A TPU is Google’s custom AI chip, offered through Google Cloud rather than as a standalone card you buy and rack yourself. A v5p pod has 8,960 chips, peak BF16 compute of 459 TFLOPS per chip, 95 GiB of HBM per chip, and 4,800 Gbps of inter-chip interconnect, and the largest schedulable v5p job Google lists is 6,144 chips.
Does Apple train Apple Intelligence on private iPhone data?
Apple says it does not. The 2024 mixture used Applebot crawls of public pages, licensed publisher sets, open-source code from license-filtered GitHub repos, math pages, and other public datasets, and the crawler honors robots.txt. The 2026 post repeats that private personal data and user interactions are left out of foundation-model training.
Where does AFM 3 Cloud Pro run, and what still runs on Apple silicon?
AFM 3 Cloud Pro runs on NVIDIA GPUs in Google Cloud under the expanded Private Cloud Compute design, with NVIDIA Confidential Computing, Intel TDX, and Google’s Titan chip in the trust stack. AFM 3 Core and AFM 3 Core Advanced run on device; AFM 3 Cloud and ADM 3 Cloud run on Apple silicon servers in Private Cloud Compute, so four of the five named 2026 models stay on Apple’s chips.
-
BUSINESS4 months agoMusk’s $914 Billion Lead Is a Public SpaceX Bet
-
NEWS2 months agoMicrosoft’s 96% Cyber Score Sits at 86.3% on the Board
-
SPORTS3 months agoFree Live Sports Streaming in 2026: What to Watch Without Cable
-
ENTERTAINMENT4 weeks agoSterling Point Holds No. 2 on Prime Video After 32 Days
-
ENTERTAINMENT3 months agoThe Odyssey’s IMAX Film Run Hit a 41-Theater Limit
-
ENTERTAINMENT3 months agoEndgame Encore’s $86 Million Trial for Infinity Vision
-
ENTERTAINMENT2 years agoAndrew Garfield’s Spider-Man Films Are No Longer Free
-
NEWS4 weeks agoG20 Cheers AI Investment After Bailey’s Frontier Cyber Warning
