← Articles
Read in another language
Technology

The AI in Your Pocket: What On-Device Models Actually Do

Jayden

Maintains the wage calculators and public-data regional information at 생활데이터랩, and analyzes technology, industry, and policy issues.

Published

Key points

  • On-device AI stopped being a demo in 2025: Apple shipped the Foundation Models framework in iOS 26, Google shipped ML Kit GenAI APIs powered by Gemini Nano, and Microsoft put Phi Silica on Copilot+ PC NPUs.
  • Every one of those shipped surfaces is task-scoped rather than an open-domain chatbot — Google's on-device APIs cover exactly four narrow jobs: summarize, proofread, rewrite, and image description.
  • Small models fit on phones by trading precision and breadth for size: Apple's on-device model uses 2 bits per weight with quantization-aware training, and Phi-3-mini takes about 1.8 GB once quantized to 4 bits.
  • A peak spec and a delivered experience answer different questions — a Copilot+ PC must carry an NPU rated at 40 or more TOPS, while on a Pixel 9 Pro Gemini Nano writes its answer at roughly 11 tokens per second.
  • The dominant real architecture is local-first with cloud fallback, not local-only: small local models take the frequent, latency-sensitive or private work, and harder requests still route to a larger cloud model.

For two years, using AI meant sending your words to someone else's computer. You typed a prompt, it traveled to a data center, a giant model answered, and the reply came back. In 2025 and 2026, a quieter shift arrived: a small but real slice of that work now happens on the device in your hand or on your desk, with nothing leaving the machine. Apple, Google, Microsoft, and Qualcomm all shipped consumer or developer surfaces built on models that run locally, and every phone and laptop launch now advertises a "neural processing unit," or NPU, as if it were the new megapixel count.

The pitch is genuinely appealing: instant responses, no network needed, no per-use cloud bill, and your data staying put. But the marketing blurs three things that are worth holding apart. There is a difference between what has actually shipped to people and what was demoed on a keynote stage. There is a difference between a benchmark or spec sheet and real-world quality on a device running off a battery. And there is a difference between a vendor's claim and independently verified behavior. Keep those three seams visible and the on-device AI story becomes much clearer — and more useful — than either the hype or the backlash suggests.

What "on-device" actually means

The models that run locally are called small language models, or SLMs, and "small" is doing real work in that name. The cloud models behind chat assistants are measured in hundreds of billions of parameters. The on-device models are measured in single-digit billions. Apple's on-device foundation model is about 3 billion parameters [source: Apple Machine Learning Research, 2025]. Microsoft's Phi Silica, which runs on Windows Copilot+ PCs, is a derivative of the Phi-3.5-mini family [source: Windows Experience Blog, 2024]. Microsoft's research prototype Phi-3-mini — the clearest public proof that this is possible at all — is 3.8 billion parameters, and when compressed it runs natively on an iPhone [source: Phi-3 Technical Report, arXiv, 2024].

Two engineering techniques make that shrinkage possible, and it helps to keep them distinct. Quantization stores a model's numbers at lower precision — dropping from 16-bit down to 4-bit or even 2-bit values — which slashes memory and speeds up math for some loss of accuracy. Apple's model uses 2 bits per weight, trained with quantization-aware training so the model learns to tolerate the coarseness [source: Apple Machine Learning Research, 2025]. Phi-3-mini, quantized to 4 bits, occupies about 1.8 GB and generates more than 12 tokens per second on an iPhone's A16 chip, fully offline [source: Phi-3 Technical Report, arXiv, 2024]. Distillation is the other technique: instead of compressing one model, you train a smaller "student" to imitate a larger "teacher," producing a genuinely new, smaller model. The SLM families shipping today are built with both.

That 1.8 GB figure is the whole story in miniature. A model that once needed a rack of server GPUs now fits in the memory budget of a phone — not by magic, but by trading away precision and breadth for a footprint that fits.

What actually shipped in 2025

This is the layer that separates on-device AI from vaporware, so it is worth being concrete about what real people and developers can actually touch today.

Apple shipped the Foundation Models framework in iOS 26 in the fall of 2025, giving developers direct access to the on-device 3-billion-parameter model. Inference is free — there is no per-call cloud bill — it works offline, and the data stays on the device [source: Apple Machine Learning Research, 2025]. Google shipped its ML Kit GenAI APIs, powered by Gemini Nano, to Android developers in May 2025. They run through the system's AICore, and input, inference, and output are all processed locally with no internet requirement and no per-call cost [source: Android Developers Blog, 2025]. Microsoft put Phi Silica on the NPUs of Copilot+ PCs, where it drives features such as Click to Do and on-device rewrite and summarize inside Word and Outlook, with a developer API following in January 2025 [source: Windows Experience Blog, 2024]. And Qualcomm announced the Snapdragon 8 Elite Gen 5 in September 2025 on a 3-nanometer process, with a Hexagon NPU it says is about 37% faster at AI than the previous generation [source: GSMArena, 2025].

Notice what these shipped products have in common: they are task-scoped, not open-ended chatbots. Google's on-device APIs are explicitly four narrow jobs — summarize, proofread, rewrite, and describe an image [source: Android Developers Blog, 2025]. Apple's model is offered as a building block for app features, not as a general assistant. This is the shipped reality, and it is deliberately narrower than the keynote impression of "AI everywhere, locally."

Benchmarks and spec sheets versus a device on battery

Here the marketing and the measured experience pull apart, and the gap is the most important thing a reader can understand.

Start with the NPU. Every Copilot+ PC must have an NPU rated at 40 or more TOPS — trillions of operations per second [source: Windows Experience Blog, 2024]. That is a peak-throughput spec, and it sounds enormous. But peak operations are not delivered words. On the same class of hardware, Phi Silica produces its first token in about 230 milliseconds and then generates up to about 20 tokens per second [source: Windows Experience Blog, 2024]. On a Pixel 9 Pro, Gemini Nano reads input at roughly 510 tokens per second but writes its answer at about 11 tokens per second [source: Android Developers Blog, 2025]. Those output speeds are perfectly usable for a two-sentence summary and visibly slow for anything long. A 40-plus-TOPS spec and an 11-tokens-per-second reality are both true; they are just answering different questions.

The same gap shows up in quality scores. Phi-3-mini posts 69% on MMLU, a standard knowledge benchmark, which its authors frame as rivaling much larger models [source: Phi-3 Technical Report, arXiv, 2024]. That is a real achievement, but a benchmark average is not open-domain reliability. The most honest statement in this entire field comes from Apple, which says plainly that its on-device model "is not designed to be a chatbot for general world knowledge" [source: Apple Machine Learning Research, 2025]. Read that again: the company shipping the model tells you not to treat it as a know-it-all. That is the benchmark-versus-reality gap stated by the vendor itself.

Battery and heat sit underneath all of this. The reason on-device inference is viable at all is that NPUs are far more power-efficient than CPUs for this work — Microsoft measured Phi Silica using about 56% less power than the same job on the CPU [source: Windows Experience Blog, 2024]. That efficiency is exactly why a phone can run a model without draining in minutes. But efficient is not free: sustained text generation still consumes energy and produces heat, which is one practical reason vendors cap what these models will attempt.

Those caps are explicit. Google states that its on-device summarizer works best on input under about 4,000 tokens, and that proofreading and rewriting are meant for short text under 256 tokens [source: Android Developers Blog, 2025]. Those are not arbitrary — they are the honest edges of where a small model on a battery delivers good results.

Claims versus what is verified

The third seam is the softest, because it is where verifiable engineering shades into marketing adjectives. "Your data stays on device" is the headline privacy claim, and for genuinely local inference it follows from the architecture — if the model runs on the phone and the network is off, the words are not leaving [source: Android Developers Blog, 2025]. That is the strongest of the claims because you can reason about it from how the system is built.

Others deserve more caution. Qualcomm describes the Snapdragon 8 Elite Gen 5 as enabling "agentic AI" with continuous on-device learning while "user data stays on device" [source: GSMArena, 2025]. "Rivals GPT-3.5" is a benchmark-anchored claim about a specific test set, not a promise about your particular question [source: Phi-3 Technical Report, arXiv, 2024]. The reliable core in each case is the part you can inspect — where the model runs, how it was quantized, how many tokens per second it measurably produces — not the adjective wrapped around it. A good habit: trust the architecture, discount the marketing.

The hybrid reality: local-first, not local-only

The cleanest way to misunderstand this shift is to imagine the cloud going away. It is not. The dominant real architecture is tiered: a small local model handles common, latency-sensitive, or private tasks, and harder requests fall back to a large model in the cloud.

Apple's shipped design is the clearest example. The on-device 3-billion-parameter model handles what it can, and heavier requests route to a larger server model — a Parallel-Track Mixture-of-Experts architecture — running on Apple silicon in what the company calls Private Cloud Compute [source: Apple Machine Learning Research, 2025]. Google and Microsoft follow the same shape: scoped tasks run locally through Gemini Nano or Phi Silica, while the heavy lifting still goes to cloud Gemini or Copilot. So when a 2026 phone advertises "on-device AI," the honest translation is usually local-first with cloud fallback, not everything on the phone. That is not a bait-and-switch; it is the sensible engineering answer to the fact that a 3-billion-parameter model and a frontier cloud model are good at different things.

What small local models can and cannot do

Strip away the layers and a practical picture remains. What on-device SLMs do well is a bounded, valuable set of jobs: summarizing a document, proofreading and rewriting text, describing an image for accessibility, extracting structure, routing a request, and calling a tool — all with low latency, offline, and without sending your data anywhere. Those are real features shipping in real products right now, and for that category the local model is often the better choice than a round trip to the cloud.

What they cannot reliably do is equally important: they are not general world-knowledge chatbots — Apple says so outright [source: Apple Machine Learning Research, 2025] — they struggle with long-context reasoning, and they do not match frontier cloud models on open-domain accuracy. Asking a 3-billion-parameter phone model to be a substitute for a hundreds-of-billions-parameter cloud assistant is asking the wrong thing of it.

What to watch

The useful posture toward on-device AI is neither the "everything is local now" hype nor the "it is just a gimmick" backlash. Something real shipped in 2025, and it is genuinely useful within its scope.

Three questions cut through the noise. First, is a capability shipping in a released OS, SDK, or chip, or is it a demo and a roadmap — because the gap between the two is where disappointment lives. Second, does a number describe a peak spec (TOPS, a benchmark average) or a delivered experience (tokens per second on a real device, within the vendor's own task caps) — because a 40-TOPS NPU and an 11-token-per-second answer are both honest and describe different things. And third, is a claim about architecture you can inspect — where the model runs, how it is quantized — or a marketing adjective like "agentic" or "rivals GPT-4." Judge on-device AI by what it verifiably does within its limits, and the technology in your pocket turns out to be smaller than the ads suggest and more useful than the skeptics allow.

Timeline

  1. Microsoft researchers publish the Phi-3 Technical Report — a 3.8-billion-parameter model that, quantized to 4 bits, occupies about 1.8 GB and runs fully offline on an iPhone A16.

    Phi-3 Technical Report (arXiv) (opens in a new tab)
  2. Microsoft details Phi Silica, a model derived from the Phi-3.5-mini family that runs on the NPU of Copilot+ PCs — machines that must carry an NPU rated at 40 or more TOPS.

    Windows Experience Blog (opens in a new tab)
  3. The Phi Silica developer API follows, opening the on-device model to Windows application developers alongside features such as Click to Do and rewrite/summarize in Word and Outlook.

    Windows Experience Blog (opens in a new tab)
  4. Google ships ML Kit GenAI APIs, powered by Gemini Nano, to Android developers — input, inference, and output all processed locally through the system's AICore, with no internet requirement and no per-call cost.

    Android Developers Blog (opens in a new tab)
  5. Apple publishes its foundation-model update: an on-device model of about 3 billion parameters at 2 bits per weight, a PT-MoE server model in Private Cloud Compute, and the plain statement that the on-device model is not designed to be a chatbot for general world knowledge.

    Apple Machine Learning Research (opens in a new tab)
  6. The Foundation Models framework reaches general availability in iOS 26, giving developers direct access to the on-device model — inference is free, works offline, and the data stays on the device.

    Apple Machine Learning Research (opens in a new tab)
  7. Qualcomm announces the Snapdragon 8 Elite Gen 5 on a 3-nanometer process, with a Hexagon NPU the company says is about 37% faster at AI than the previous generation.

    GSMArena (opens in a new tab)

Analysis

Shipping is the first filter, and it is a strict one

The gap where disappointment lives is between what was demonstrated on a keynote stage and what arrived in a released OS, SDK, or chip. Apple's Foundation Models framework in iOS 26, Google's ML Kit GenAI APIs, Microsoft's Phi Silica on Copilot+ PCs and Qualcomm's Snapdragon 8 Elite Gen 5 all clear that bar. Applying the filter first removes most of the argument about on-device AI before it starts.

Small is not a marketing adjective — it is an order of magnitude

Cloud chat models are counted in hundreds of billions of parameters; on-device models are counted in single-digit billions. Apple's runs at about 3 billion parameters, and Phi-3-mini — the clearest public evidence that any of this is possible — is 3.8 billion. That difference in scale is what makes free, offline, per-call-cost-free inference possible, and it is also what makes open-domain reliability impossible.

Peak operations are not delivered words

A Copilot+ PC must clear 40 or more TOPS, which sounds enormous. On the same class of hardware Phi Silica returns a first token in about 230 milliseconds and then generates up to about 20 tokens per second; on a Pixel 9 Pro, Gemini Nano writes at about 11 tokens per second. Both numbers are honest. They simply measure different things — the ceiling of the silicon versus the rate a user actually reads.

The vendors' own caps are the most trustworthy numbers they publish

Google states that its on-device summarizer works best under about 4,000 tokens of input and that proofreading and rewriting are meant for text under 256 tokens. Apple states outright that its on-device model is not designed to be a chatbot for general world knowledge. These are limits published by the companies selling the feature — which makes them a firmer guide to real capability than any benchmark average.

Local-first with cloud fallback, not local-only

The tiered design is the point, not a compromise. A small local model handles the frequent, latency-sensitive or private work; harder requests route to a much larger model in the cloud — Apple's PT-MoE server model in Private Cloud Compute, or cloud Gemini and Copilot. Reading a 2026 'on-device AI' claim as 'everything happens on the phone' is the single most common misunderstanding of the shift.

Comparison

Layer A — what actually shipped, and how narrow its scope is (vendor announcements, 2024-2025)
VendorShipped surfaceTask scope as stated
AppleFoundation Models framework, generally available in iOS 26 (fall 2025)A component for building app features on an on-device model of about 3 billion parameters — explicitly not a general assistant
GoogleML Kit GenAI APIs powered by Gemini Nano, to Android developers (May 2025)Exactly four narrow jobs: summarize, proofread, rewrite, image description
MicrosoftPhi Silica on Copilot+ PC NPUs; developer API from January 2025Click to Do, and on-device rewrite/summarize inside Word and Outlook
QualcommSnapdragon 8 Elite Gen 5, 3-nanometer process (September 2025)Silicon, not a feature: a Hexagon NPU the company says is about 37% faster at AI than the previous generation
Layer B — peak spec versus delivered experience. Figures are quoted with the hedges their sources used; they come from different producers, devices, and measurement bases, so they are not directly comparable to one another.
Figure as publishedWhat kind of number it isWhat it does not tell you
40 or more TOPS — the NPU every Copilot+ PC must have (Windows Experience Blog, 2024)A certification floor and a peak-throughput specHow many words per second any particular model actually produces
About 230 milliseconds to first token, Phi Silica (Windows Experience Blog, 2024)A measured latency on that class of hardwareHow long the rest of a long answer will take
Up to about 20 tokens per second, Phi Silica (Windows Experience Blog, 2024)A ceiling — the phrasing is 'up to'The rate under sustained load, heat, or a low battery
Roughly 510 tokens per second reading input, about 11 tokens per second writing output, Gemini Nano on a Pixel 9 Pro (Android Developers Blog, 2025)Measured on one device; reading and writing are separate operationsWhether another phone or another task reaches the same rate
More than 12 tokens per second, Phi-3-mini on an iPhone A16, fully offline (Phi-3 Technical Report, 2024)A floor from a research model, not a shipped consumer featureWhere the actual rate lands above that floor
69% on MMLU, Phi-3-mini (Phi-3 Technical Report, 2024)A benchmark average on a standard knowledge testReliability outside the test set, in open domains
Layer C — claim versus what can be independently checked
ClaimKind of evidenceHow far it can be verified
"Your data stays on the device" (Android Developers Blog, 2025)Follows from the architecture: the model runs locally through AICore and needs no internetFirm — it can be reasoned from how the system is built, and tested with the network off
"Agentic AI" with continuous on-device learning (Qualcomm, via GSMArena, 2025)A vendor product description at launchSoft — the wording is a marketing frame, not a specification anyone can measure
"Rivals GPT-3.5" (Phi-3 Technical Report, 2024)A benchmark-anchored claim about a specific test setBounded — true of that test set, not a promise about a particular question you ask
"Not designed to be a chatbot for general world knowledge" (Apple Machine Learning Research, 2025)A limitation stated by the company that shipped the modelStrongest of the four — the seller is naming the boundary rather than the capability

Process

  1. Check the shipping surface

    Is the feature in a released OS, SDK, or chip you can buy today — or is it a keynote demo and a roadmap?

  2. Read the stated task scope

    Shipped on-device features are narrow by design. Google names four jobs; Apple offers a building block, not an assistant.

  3. Separate the peak spec from the delivered rate

    TOPS and benchmark averages describe a ceiling. Tokens per second on a real device, under the vendor's own caps, describes the experience.

  4. Look up the vendor's own caps

    Input limits such as under about 4,000 tokens for summarizing, or under 256 for rewriting, mark where a battery-powered model still does good work.

  5. Ask where the model runs and how it was compressed

    Where inference happens, and at how many bits per weight, are inspectable facts — and they are what a privacy claim actually rests on.

  6. Discount the adjectives

    'Agentic', 'rivals GPT-3.5', and similar wrappers are claims about framing or a test set. Judge the parts you can inspect instead.

Sources

  1. Apple Machine Learning Research — Updates to Apple's On-Device and Server Foundation Language Models (2025-06-09).View source (opens in a new tab)
  2. Android Developers Blog — On-device GenAI features with ML Kit and Gemini Nano (2025-05-20).View source (opens in a new tab)
  3. Windows Experience Blog — Phi Silica, small but mighty on-device SLM (2024-12-06).View source (opens in a new tab)
  4. Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone (arXiv:2404.14219, 2024).View source (opens in a new tab)
  5. GSMArena — Qualcomm Snapdragon 8 Elite Gen 5 features and specs (2025-09-25).View source (opens in a new tab)

Tags

  • #on-device-ai
  • #small-language-models
  • #edge-ai
  • #npu
  • #quantization
  • #ai-privacy