Apple is currently engaging in discussions with a small startup based in Silicon Valley that claims to have the technology to compress advanced artificial intelligence models enough to function directly on iPhones, as stated by the startup's CEO in an interview with CNBC. The company, PrismML, emerged as a spinout from the California Institute of Technology and is supported by Khosla Ventures. Recently, PrismML unveiled a compressed version of Alibaba's open-source Qwen model, which has been reduced from approximately 54 GB to under 4 GB, enabling all its 27 billion parameters to run on iPhone 15 and later models.
Babak Hassibi, PrismML's CEO, informed CNBC that Apple, along with other firms, is currently assessing the performance, speed, and energy efficiency of the startup’s models on various devices. He described the ongoing conversations as preliminary and uncertain in terms of outcomes, though he noted positive progress.
This release coincides with Apple’s recent launch of the public beta for iOS 17, which is intended to give iPhone users access to substantial updates to Siri. The objective is to enhance Siri’s competitiveness with AI assistants from OpenAI and Anthropic, all while processing more personal information directly on devices. This strategy could potentially overcome a significant limitation of Apple’s AI initiatives, as high-performing models generally demand excessive memory and processing power not typically available on smartphones.
While Apple can utilize cloud-based models for complex requests, running AI directly on iPhones offers reduced latency by eliminating the reliance on remote servers. This shift could also decrease cloud computing expenses and strengthen the company's focus on privacy, as certain features could function without requiring internet access.
Carolina Milanesi, president and principal analyst at Creative Strategies, emphasized that smaller AI models could enable Apple to incorporate more advanced features directly onto iPhones, including those requiring sensitive personal health data, computational photography, and video generation. The startup claims to achieve this model compression by radically simplifying how data is stored internally, reducing memory requirements drastically. They compare their technology's advancements to the transition from eight-bit to four-bit computing, achieving up to 15 times reduction in memory, accelerated response times, and decreased energy consumption compared to traditional models.
However, Hassibi acknowledged that there are trade-offs, with their models showing a slight reduction in overall performance, particularly in factual recall compared to reasoning, mathematics, and coding capabilities. PrismML is making two compressed versions of its model available for free, targeting everyday devices like iPhones and MacBooks.
The technology, which originated from Hassibi's research group at Caltech, is exclusively licensed to PrismML. The company successfully raised $16.25 million in seed funding in March, backed by Khosla Ventures and other investors. Looking ahead, Hassibi mentioned plans to work on Google's open-source Gemma model, followed by larger models typically reliant on data center hardware.
PrismML’s technology has the potential to extend into domains beyond smartphones and laptops, including robotics and autonomous systems requiring swift decision-making without cloud reliance. "Local intelligence that operates quickly is crucial," Hassibi stated.
Currently, Apple already employs portions of its AI capabilities locally, such as translation and other features tightly linked to users' personal data. More demanding tasks are typically sent to Apple’s private cloud infrastructure or external models. Horace Dediu, founder of Asymco, suggested that Apple likely aims to maintain most routine Siri interactions on-device while relegating more intensive tasks to the cloud.
The benefits extend beyond memory efficiency; Apple is exploring how substantial models can be effectively integrated within the physical constraints of its devices. Dediu pointed out that by keeping the majority of common requests local, Apple could offer improved latency, enhanced privacy, and potentially reduced costs associated with cloud services. Apple's control over both the iPhone's hardware and software could give it an edge in maximizing the efficiency of AI in its devices.
Industry analysts cautioned, however, that claims made by PrismML need validation outside experimental settings. Performance in more extensive tasks, battery life during multitasking, and reliability across varying scenarios will be pivotal. Tarun Pathak, research director at Counterpoint Research, highlighted the need for thorough testing across millions of queries and diverse device types.
PrismML's launch occurs amid significant discussions regarding AI efficiency and its implications for the demand for memory chips and the expensive infrastructure of data centers. Memory constraints in consumer electronics and AI setups have been escalating, leading Morgan Stanley to predict that Apple's costs for dynamic random access memory could surge by 190% year-on-year by fiscal 2027. Amidst this, PrismML's technology suggests that cloud models requiring eight GPUs could potentially run on just one, facilitating the transition of models from servers to personal devices.
However, analysts, like Gil Luria from D.A. Davidson, argue that while smaller models may shift some processing from data centers to phones, the need for processors and memory will persist. "You're still going to need the GPU, and you're still going to need the memory," Luria stated, pointing out that handling AI on individual devices might sometimes be less energy-efficient compared to centralized cloud infrastructure.
Despite the potential for improved AI efficiency, investors have reacted cautiously to signals that AI could require less memory, as evidenced by the sharp decline in Micron shares following Google's revelations about memory optimization without compromising performance. With PrismML now providing its models for public scrutiny, both users and investors can assess the viability of its promises outside controlled environments. For Apple, integrating more advanced AI capabilities directly on its devices could bolster Siri's functionality while maintaining the levels of privacy and hardware integration that differentiate its products. "A combination of cloud and on-device AI can deliver a more comprehensive and privacy-friendly AI experience," concluded Counterpoint's Pathak.



