
Agentic AI · Privacy 2026
On-device AI vs cloud — keep inference where the data is allowed to be
The prompt is the data
If a user types a name, a contract clause, or a medical note into a box that calls a model in another country, that text has left the device. Logs, embeddings, and human-review queues often leave with it. Calling the feature “AI” does not change the geography.
On-device inference keeps that text on the handset or the laptop. The trade-off is capability: a small local model will not match a frontier API on open-ended reasoning. We treat that as a product choice, not a branding choice.
What we actually run on-device
My Instincts is our privacy-first daily brief. The pipeline is deterministic — extract, synthesise, judge significance — and a small on-device model is used only for final phrasing. Android is live. iOS is in progress. There is no Play Store URL to invent, and nothing leaves the phone for inference.
That architecture is the opposite of “paste the day into ChatGPT.” It is also slower to ship, and it is the right shape when the brief is a private journal, not a public knowledge base.
How we choose a tier
Engineering tiers — not a legal determination.
| Tier | When we use it |
|---|---|
| On-device / on-prem | The brief says personal or client data must not leave hardware the user or the company operates. Daily briefs, device-local assistants, air-gapped ops. |
| In-region managed | Personal data is in play, but a UAE-region endpoint and a written processing path are acceptable to the owner. We still minimise what goes in the prompt. |
| Public cloud API | The data class is not personal, or counsel has a transfer basis the owner accepts. Useful for bursty, high-capability tasks that do not need to live on a phone. |
Hybrid is normal
Retrieve and judge locally. Draft with a stronger model only on text that has already been stripped. Keep the write-back behind a human. That pattern shows up in both our on-device work and our agent builds.
If a vendor tells you every workload must be on-prem, or that a consumer chatbot is fine for client files, ask them to put the data path on one page. We will put ours on the statement of work.
FAQ
Is on-device always more private?
It removes the network hop for inference. It does not remove on-device backups, screenshots, or a later sync the product might add. Privacy is the whole path, not the model card.
Can an agent run entirely on-device?
A narrow agent can — if its tools are local. The moment it must call your CRM or send mail, you are back to credentials, logs, and a boundary. We design that boundary explicitly.
Do you fine-tune Arabic models on-device?
We use small on-device models for phrasing, not as a claim that we operate a sovereign LLM lab. Arabic and English as product languages are a separate, scoped piece of work.
