Apple ships a language model, an image generator and a room scanner on the device. What each can do from Flutter, and the honest limits of all three.
Every "add AI to your Flutter app" tutorial ends the same way: get an API key,
add a billing card, send the user's text to somebody else's server, and hope the
latency is tolerable.
On iOS 26 and macOS 26 there is a different option that most Flutter developers
have not tried, because the plugins did not exist. Apple ships a language model,
an image generator and a LiDAR room scanner on the device. No key, no
account, no per-token cost, no network — and nothing the user types leaves their
phone.
I built three packages to reach them from Flutter. This is what each one can
genuinely do, and where each one stops.
final session = await LanguageModelSession.create( instructions: 'You summarise text in one sentence.',);print(await session.respond(article));await session.dispose();
That runs entirely on the device. It streams, it supports tool calling, and —
the part that matters most in practice — it does **schema-constrained structured
output**. You describe the shape you want and the model is constrained to produce
it, rather than being asked nicely in a prompt and then parsed hopefully.
If you have ever written a retry loop around jsonDecode because a hosted model
returned prose instead of JSON, that is the problem this removes at the source.
It is worth dwelling on this one, because it is the difference between a demo
and something you can put in a product.
The usual pattern with a hosted model is: write a prompt asking politely for
JSON, get back something that is usually JSON, wrap jsonDecode in a
try/catch, add a retry, and accept a failure rate you cannot drive to zero. Every
production LLM integration has this scar tissue.
Schema-constrained generation removes the problem rather than mitigating it. The
model is constrained during sampling to tokens that keep the output valid
against your schema. It is not being asked to produce the right shape; it is
unable to produce a wrong one.
For anything where the model's output feeds code rather than a human — extract
these fields, classify into these categories, pull the dates out of this text —
that is the difference between a feature you ship and a feature you keep
apologising for.
The model can also call functions you expose. The pattern is the familiar one —
you describe the tools, the model decides when to call them, you run them and
hand back results — but with the round trip happening entirely on the device.
The practical use is less "agents" and more "give the model access to the user's
own data without that data going anywhere". A tool that reads the user's local
notes or transactions lets the model answer questions about them, while the notes
themselves never leave the phone. That is not achievable with a hosted model at
any price.
Where it stops. This is a small model. It is very good at summarising,
extracting, classifying, rewriting and answering questions about text you give
it. It is not a frontier model and it will not reason its way through a hard
novel problem. Treat it as a fast, free, private text-processing engine and it
is excellent. Treat it as a substitute for a large hosted model and it will
disappoint you.
apple_foundation_models — iOS and macOS.
await for (final image in ImageCreator.generate(prompt)) { setState(() => _preview = image);}Image Playground generates on the device too, and the API streams **partial
images as they are made**. That is not a nicety, it is the entire user
experience. Here is the timing I measured:
| | |
| --- | --- |
| first image | 6.8s |
| second image | 10.3s |
| size | ~4 MB PNG each |
Seven to ten seconds of a spinner feels broken. Seven to ten seconds of an image
visibly resolving in front of you feels like something is being made for you.
Same wait, completely different experience — which is why the package streams by
default rather than offering it as an option.
The same package also wraps the native text field that carries Writing Tools
and Genmoji, so users get the system rewrite, proofread and summarise
affordances they already know from Notes and Mail, in your app, without you
building any of it.
Look at the timings again and think about what the alternative looks like.
A modal spinner for 6.8 seconds is a long time. Long enough that a meaningful
share of users will assume it has hung and back out. Long enough that on the
second image, at 10.3 seconds, some will kill the app.
The same 10.3 seconds spent watching an image emerge from noise into shape is not
a wait at all — it is the product. The user is watching their idea being made.
Nobody backs out of that, because there is visible progress on every frame.
This is the general lesson for on-device generation, not a quirk of this API. The
compute is slower than a datacentre GPU. You cannot fix that. What you can do is
make the latency visible and interesting rather than hidden behind a spinner,
and the perceived experience inverts.
Budget around 4 MB per PNG, which matters if you are storing or syncing them.
apple_intelligence — iOS and macOS.
final controller = RoomScanController();
RoomPlan uses the LiDAR sensor to scan a room and hand back structured geometry:
| | |
| --- | --- |
| walls, floors | surfaces with metre dimensions and a transform |
| doors, windows, openings | found within the walls |
| objects | recognised furniture — chair, table, bed, storage… |
| usdzPath | a USDZ model, ready for AR Quick Look |
| raw | RoomPlan's own encoding, untouched |
Real dimensions in metres, not a point cloud you have to interpret. For anyone
building for interiors, property, removals, insurance or furniture retail, that
is the difference between a demo and a product.
The package deliberately hosts Apple's own scanning UI rather than
reimplementing it. Users have seen that interface in Apple's Measure app; the
coaching overlays that tell you to move slower or point at the wall are tuned by
people with the sensor data. Replacing that with a custom camera view would be
worse in every way that matters.
roomplan — iOS, LiDAR required.
The USDZ output gets the attention because it drops straight into AR Quick Look.
The structured geometry is the more valuable half.
Because you get walls with metre dimensions, doors and windows located within
them, and recognised furniture, you can answer questions rather than just display
a model: will this sofa fit through that door? *how much paint for these
walls?* what is the floor area, excluding the built-in wardrobe?
That is the difference between a scanning feature and a product. Estate agents,
removals firms, insurers, interior retail and trades all have a version of that
question, and all of them currently answer it with a tape measure and a phone
call.
These are Apple frameworks, so the constraints are Apple's:
design a fallback path — all three packages expose a capability check, and you
should call it before showing any UI that depends on the feature.
Availability is not the same as "the device is new enough".
of wrapping changes that.
Which is why the honest framing is not "replace your hosted model". It is: for
the substantial set of tasks these handle, you get them for free, instantly,
offline, and without any user data leaving the device — and you fall back to
whatever you use today when the device cannot help.
That fallback path is not a workaround. It is the design.
There are three separate questions and they have three different answers, which
is why a single isSupported boolean is not enough:
RoomPlan needs LiDAR.
perfectly capable device.
Only the third can change while your app is running, and it is the one people
forget. Query availability at the point of use rather than caching a result from
launch, and design the UI so an unavailable feature is absent rather than
present-and-failing. A greyed-out button with a tooltip explaining an OS
requirement is worse than not showing the button, because it advertises
something the user cannot act on.
Three reasons that survive contact with a real product.
Cost. Per-token pricing turns every active user into a running bill.
On-device inference costs nothing and does not scale with success.
Latency. No round trip. For summarise-as-you-type or classify-on-paste, the
difference between 40ms and 900ms is the difference between a feature people use
and one they wait for.
Privacy you can state simply. "Your notes never leave your phone" is a
sentence you can put in an App Store listing and defend. It is also the only
version of that promise that is true without a legal team.
To be clear about when not to reach for these:
| | On-device | Hosted |
| --- | --- | --- |
| Cost per call | none | per token |
| Latency | ~40ms to first token | 300ms–2s round trip |
| Works offline | yes | no |
| Data leaves device | no | yes |
| Reasoning depth | limited | frontier |
| Platform | iOS/macOS 26+ | anywhere |
| Availability | user must enable | always |
The right architecture for most apps is both: on-device for the high-frequency,
low-complexity, privacy-sensitive work, hosted for the occasional hard question,
and a capability check deciding which one runs. That is not a compromise, it is
just using the cheaper tool for the easier job.
All three packages are MIT licensed with public source, and all three hold a
perfect 160/160 on pub.dev. Start with whichever problem you actually have —