Editorial illustration for Apple unveils third‑gen foundation model, AFM 3 Cloud shows 36% boost
Apple unveils third‑gen foundation model, AFM 3 Cloud...
For years, Apple's AI was publicly seen as a laggard. That story ended Tuesday. The third-generation AFM 3 Cloud arrived as a wholesale correction, with internal benchmarks showing a 36% jump in response satisfaction.
It follows instructions 21% better. But the starkest leap is visual: the new model wins on 37.8% of image prompts. That's a dramatic shift from a paltry 9.6% for its 2025 predecessor.
For clients demanding the peak, a Pro variant tacks on another 10% for text and 14% for images.
We also see consistent gains in our single-sided evaluations, which score responses independently along multiple dimensions: AFM 3 Cloud delivers a roughly 36 percent relative improvement in overall response satisfaction and a 21 percent relative improvement in instruction following performance over the 2025 AFM Server model. Further, for image understanding, where the model interprets and reasons over visual inputs, AFM 3 Cloud showed significant improvement over its predecessor from last year, earning preference on 37.8 percent of prompts compared to just 9.6 percent for its 2025 baseline. Finally, AFM 3 Cloud Pro provides an even further improvement over our AFM 3 Cloud, achieving a relative improvement in overall response satisfaction of roughly 10 percent for text and 14 percent for image understanding overall.
Tripling performance in image understanding, jumping to that 37.8% preference rate, is an admission. The old architecture was limited. Now Apple signals it can compete on raw capability.
But do these lab metrics—that 36% satisfaction lift—mean a useful answer in a noisy kitchen? They must. The model is less often stupid.
That's the foundation. The race just got less predictable.
Common Questions Answered
What performance improvements does Apple's AFM 3 Cloud foundation model demonstrate?
Apple's third-generation AFM 3 Cloud shows a 36% jump in response satisfaction compared to previous versions, along with 21% better instruction following capabilities. The most dramatic improvement is in visual understanding, where the new model wins on 37.8% of image prompts, representing a significant shift in Apple's competitive positioning in AI.
How does AFM 3 Cloud's image understanding capability compare to previous Apple models?
The AFM 3 Cloud has tripled performance in image understanding, achieving a 37.8% preference rate on image prompts. This represents a dramatic improvement that signals Apple's old architecture was limited and that the company can now compete on raw capability in visual AI tasks.
What does the 36% satisfaction boost in AFM 3 Cloud mean for real-world applications?
The 36% internal benchmark satisfaction lift indicates that the model is significantly less prone to errors and provides more reliable responses in practical scenarios. This foundational improvement in reducing incorrect outputs is essential for the model to perform effectively in real-world environments like noisy kitchens, where accuracy and reliability are critical.
Why is Apple's AFM 3 Cloud release considered a turning point for the company's AI reputation?
For years, Apple's AI was publicly perceived as lagging behind competitors, but the AFM 3 Cloud arrival represents a wholesale correction to that narrative. With its 36% satisfaction boost, superior instruction following, and dominant 37.8% image prompt performance, Apple demonstrates it can now compete on raw capability and has made the AI race less predictable.
Further Reading
- Apple Intelligence Foundation Language Models Tech Report 2025 — Apple Machine Learning Research
- Updates to Apple's On-Device and Server Foundation Language Models — Apple Machine Learning Research
- Apple Updates Its On-Device and Cloud AI Models, Introduces a New Developer API — The Batch
- Apple Intelligence: Unveiling Foundation Models Powering the Future of iOS, iPadOS and macOS — Synced
- Apple Intelligence Foundation Language Models - arXiv — arXiv