Loading article…
Apple is evaluating PrismML’s Bonsai 27B model, a 27‑billion‑parameter LLM compressed to 4 GB and runnable on iPhone, iPad and Mac – see why it matters for
Apple is in early talks with startup PrismML about its Bonsai 27B model, a 27‑billion‑parameter LLM compressed to roughly 4 GB and capable of running natively on iPhone, iPad and Mac [1]. The discussions matter because they could let Apple bring server‑class AI capabilities onto its devices without relying on cloud inference, a key differentiator in the mobile AI race.
| At a glance | |
|---|---|
| Company | PrismML |
| Model | Bonsai 27B (27 B parameters) |
| Size | ~4 GB after compression |
| Release | Developer preview API launched July 14, 2026 |
PrismML announced Bonsai 27B on July 14, 2026, claiming it runs at up to 163 tokens/second in 1‑bit mode on an NVIDIA RTX 5090 and 87 tokens/second on Apple’s M5 Max chip [1]. The model’s 4 GB footprint fits within the roughly 6 GB of usable memory on a 12 GB iPhone, leaving room for KV cache and activations – a size no conventional 27 B model can achieve [1]. The startup also released the model under an Apache 2.0 license and opened a limited‑time developer preview API [1].
The Information reported that Apple has already held meetings with PrismML to explore “ways it could use its technology” [2]. Apple’s broader AI strategy has focused on shrinking large language models for on‑device use, but it has been constrained by the memory demands of server‑sized models [4]. PrismML’s ability to compress the 54 GB Qwen 3.6 model to 4 GB without performance loss directly addresses this bottleneck [4]. If Apple adopts the tech, it could enable more sophisticated on‑device AI features—such as advanced chat, reasoning, and autonomous agents—than the current generation of phone‑friendly models, which typically run only a few billion parameters [2][4].
Apple’s rivals, notably Google and Microsoft, are investing heavily in cloud‑centric AI services, while Apple has emphasized privacy‑first, on‑device processing. Securing a method to run a 27 B model locally would give Apple a unique value proposition: high‑quality AI without transmitting data to the cloud. The move also aligns with Apple’s recent collaborations, such as co‑developing a new Siri AI with Google for iOS 27 [2]. However, the talks remain “very early,” and no formal partnership has been confirmed [1].
If Apple successfully integrates PrismML’s compression tech, it could reshape the mobile AI landscape by delivering server‑grade capabilities on consumer devices, but the outcome hinges on technical validation and strategic alignment in the coming months.
Coverage is mostly measured — 286 of 289 reports stay neutral.
Every Monday — the token unlocks, Fed dates & catalysts set to move crypto and markets this week. So you’re never blindsided.
Free · 3-min read · one-click unsubscribe
AI-assisted synthesis by the TrendWatcher Editorial Desk · sourced from 4 outlets · Jul 14, 2026 · How we report
Apple has announced an event scheduled for Wednesday, September 9, 2026, at 10 a.m. Pacific Time.
The streaming service now costs $14.99 per month or $119 per year.
John Ternus is preparing to take over as the CEO of Apple.
Deliveries and in-store availability for the new Mac mini and Mac Studio models are scheduled for September 22, 2026.