Skip to content

Agentic AI · Privacy 2026

On-device AI vs cloud — keep inference where the data is allowed to be

A cloud API is a transfer of the prompt. An on-device model is a product decision. We have shipped the second on Android; we use the first when the brief allows it.

The prompt is the data

If a user types a name, a contract clause, or a medical note into a box that calls a model in another country, that text has left the device. Logs, embeddings, and human-review queues often leave with it. Calling the feature “AI” does not change the geography.

On-device inference keeps that text on the handset or the laptop. The trade-off is capability: a small local model will not match a frontier API on open-ended reasoning. We treat that as a product choice, not a branding choice.

What we actually run on-device

My Instincts is our privacy-first daily brief. The pipeline is deterministic — extract, synthesise, judge significance — and a small on-device model is used only for final phrasing. Android is live. iOS is in progress. There is no Play Store URL to invent, and nothing leaves the phone for inference.

That architecture is the opposite of “paste the day into ChatGPT.” It is also slower to ship, and it is the right shape when the brief is a private journal, not a public knowledge base.

How we choose a tier

Engineering tiers — not a legal determination.

TierWhen we use it
On-device / on-premThe brief says personal or client data must not leave hardware the user or the company operates. Daily briefs, device-local assistants, air-gapped ops.
In-region managedPersonal data is in play, but a UAE-region endpoint and a written processing path are acceptable to the owner. We still minimise what goes in the prompt.
Public cloud APIThe data class is not personal, or counsel has a transfer basis the owner accepts. Useful for bursty, high-capability tasks that do not need to live on a phone.

Hybrid is normal

Retrieve and judge locally. Draft with a stronger model only on text that has already been stripped. Keep the write-back behind a human. That pattern shows up in both our on-device work and our agent builds.

If a vendor tells you every workload must be on-prem, or that a consumer chatbot is fine for client files, ask them to put the data path on one page. We will put ours on the statement of work.

FAQ

Is on-device always more private?

It removes the network hop for inference. It does not remove on-device backups, screenshots, or a later sync the product might add. Privacy is the whole path, not the model card.

Can an agent run entirely on-device?

A narrow agent can — if its tools are local. The moment it must call your CRM or send mail, you are back to credentials, logs, and a boundary. We design that boundary explicitly.

Do you fine-tune Arabic models on-device?

We use small on-device models for phrasing, not as a claim that we operate a sovereign LLM lab. Arabic and English as product languages are a separate, scoped piece of work.

Tell us what you want the system to do

We will say where a model helps, where code is the better answer, and whether a build is worth starting.