Apple Exploring Ways to Run Much Larger AI Models Directly on iPhones - MacRumorsOpen MenuShow RoundupsShow Forums menuVisit ForumsOpen Sidebar
Skip to Content

Apple Exploring Ways to Run Much Larger AI Models Directly on iPhones

Apple has held meetings with PrismML about ways it could use the startup's technology to run much larger AI models directly on iPhones, according to The Information.

ios 27 siri animation
The report said PrismML has managed to shrink down Alibaba's open-source large language model Qwen 3.6 to run entirely on an iPhone 17 Pro. The model has 27 billion parameters, which is larger than Apple's on-device AFM 3 Core Advanced model with 20 billion parameters. Apple's model powers iOS 27 enhancements such as Siri AI's more expressive voices and improved systemwide dictation on iPhone 17 Pro and iPhone Air models.

Unlike with AFM 3 Core Advanced, all of Qwen 3.6's parameters can be active at the same time.

"One new on-device Apple model has 20 billion parameters but uses a so-called sparse architecture, in which only 1 billion to 4 billion parameters are active at a time," the report said, in reference to AFM 3 Core Advanced. "In the case of PrismML's on-device model, all 27 billion parameters are active at the same time."

Larger models running directly on iPhones would allow for more Apple Intelligence features to run on device instead of on Apple's Private Cloud Compute servers, which could reduce Apple's costs and further enhance user privacy.

Popular Stories

Mac mini vs Studio Feature Sans Text 1

Apple Caught Off Guard by AI Demand for Mac Mini and Mac Studio

Sunday August 30, 2026 8:52 am PDT by
Apple's unusually timed announcement of new Mac mini and Mac Studio models this week was driven by unexpectedly strong enterprise appetite for AI hardware, according to The Information. Apple normally releases new Mac models in the autumn, closer to October or November, making this week's announcement unusually early, falling just before the anticipated arrival of new iPhone models. The...
apple watch series 12

Three All-New 'Audio Intelligence' Features Arrive With Apple Watch Series 12 and Ultra 4

Wednesday September 9, 2026 2:23 pm PDT by
Apple today announced "Audio Intelligence," a new suite of Apple Intelligence-powered features for the Apple Watch Series 12 and Apple Watch Ultra 4 that use the built-in microphone to help users catch missed moments in conversation and stay aware of important sounds. The Apple Watch's Audio Intelligence comprises four features: Siri Recap: Uses ambient listening to generate high-level...
iOS 27 Next to iPhone

Here's When iOS 27 Rolls Out Today in Every Time Zone [Update: It's Out]

Sunday September 13, 2026 3:00 am PDT by
Update 10:04 a.m.: iOS 27 is rolling out now, though it may take a bit for all users to see it, so keep checking! Apple is about to release iOS 27, which will finally deliver more advanced Siri AI capabilities as well as a variety of other refinements, improvements, and new features to iPhones. It's Apple's biggest software update of the year, and Apple announced at Wednesday's iPhone event...

Top Rated Comments

Taq'aix Avatar
10 weeks ago
The AI bubble can’t burst soon enough.
Score: 27 Votes (Like | Disagree)
smeagol Avatar
10 weeks ago
Ultimately, Apple shot themselves in the foot by either being stingy with RAM across all devices for decades, or by upgrading to higher capacity memory prohibitively expensive in the name of profits, saying nonsense like 8GB on an Apple device is like 16GB for everyone else. AI came along and told the truth, 8GB is 8GB.
Score: 23 Votes (Like | Disagree)
10 weeks ago


The report said PrismML has managed to shrink down Alibaba's open-source large language model Qwen 3.6 to run entirely on an iPhone 17 Pro. The model has 27 billion parameters,
There are many comments here, and not one about PrismML's new technology.

What they have done is invent a new way to compress a neural network to one bit per parameter. This means each parameter is just a one or a zero. Not only does this save space, it saves a LOT of space. Now Apple's 10B-parameter on-device model will fit in just over 1GB of RAM and hence comfortably into a 6 GB iPhone. (The iPhone 15 has only 6GB of RAM.)

Not only does it save space, but it also runs with less energy because it is very easy to multiply by 1 or by 0. Most of us can do that kind of math in our heads.

How does it work exactly? I don't know yet. I assume it is not so easy as simply normalizing all values to the 0...1 range and thresholding at 0.5. I suspect that replicating the "important" parameters is involved, but I don't know how you would find them.

PrismML says they are not done yet. Of course, a width of 1 is the shortest possible, but maybe they are reducing the number of parameters without doing much harm?

PrismML says the work is based on mathematics. They don't claim AI breakthroughs or better code. This might mean they have some Linear Algebra experts.

Maybe someone here has some better insight?
Score: 20 Votes (Like | Disagree)
DanteHicks79 Avatar
10 weeks ago
👏 NOBODY 👏 WANTS 👏 THIS 👏 AI 👏 GARBAGE 👏
Score: 20 Votes (Like | Disagree)
turbineseaplane Avatar
10 weeks ago

Speak for yourself. If nobody wanted it, it wouldn’t exist.
AI does not exist right now, in its current form, because of demand for it.
Score: 16 Votes (Like | Disagree)
10 weeks ago
This is the future. If we can have current model performance on-device, that will help solve a lot of the energy problems. It's likely years away (if it ever gets there), but it should be one of the goals.
Score: 16 Votes (Like | Disagree)

🔗 Related Apple News & Rumors

Stay updated with the latest Apple ecosystem news and verified rumors