Guides8 min read

Sovereign AI in the UAE: How Banks Can Deploy Private LLMs and Keep Data in the Country

What sovereign AI means for a UAE bank, why regulators and data rules push LLM workloads onshore, which open-weight models are worth testing, and a reference architecture for running a private LLM on in-country GPUs.

Most banks in the UAE now have a list of generative AI use cases. Credit memos, KYC file review, policy search, complaint triage. What slows them down is the question every security and compliance team asks first: where do the prompts and documents go, and who can see them?

For many institutions the answer that gets approved is a private LLM: an open-weight model running on GPUs inside the bank's data center or an in-country facility, with no external model API in the request path. This guide explains what that setup involves, why the regulatory picture in the UAE points toward it, and how to make it fast and cheap enough to use every day.

What is sovereign AI?

Sovereign AI means the three parts of an AI system sit under one jurisdiction and under your control:

  • The model. You hold the weights, choose the version and decide when it changes. No provider can deprecate it or alter its behavior without you knowing.
  • The data. Prompts, retrieved documents, outputs and logs stay inside your perimeter and inside the country.
  • The hardware. Inference runs on GPUs you own or rent in-country, under contracts and access controls your institution can audit.

A private LLM is the practical form of this. It is usually an open-weight model deployed on-premise or in a UAE data center, connected to internal systems through the bank's own identity and logging stack.

Why do UAE banks care about data residency for LLMs?

Three things tend to come up in every review.

Central Bank expectations on enabling technologies

In 2021 the Central Bank of the UAE, together with the SCA, DFSA and FSRA, issued guidelines for financial institutions adopting enabling technologies, covering AI, big data analytics and cloud computing among others. For cloud arrangements, the guidelines generally expect institutions to keep contractual rights of audit and access (for themselves and for their supervisor) and to maintain a documented exit plan for each arrangement. In 2026 the Central Bank followed with guidance on the responsible adoption of AI and machine learning, which makes clear that accountability stays with the licensed institution when an AI system is outsourced to a vendor.

A hosted model API is hard to fit into that frame. Audit rights over a foundation model provider's infrastructure are difficult to negotiate, and an exit plan for a proprietary model you cannot download is mostly a migration project you have not started yet.

The PDPL and cross-border transfers

Federal Decree-Law No. 45 of 2021 on the Protection of Personal Data (the PDPL) restricts transfers of personal data outside the UAE. Transfers are generally permitted to countries with an adequate level of protection, or under specific conditions set out in the law. Banking and credit data that is covered by its own sector legislation falls partly outside the PDPL, but that does not loosen things in practice: sector rules from the Central Bank often set stricter expectations on where customer data is stored and processed. Free zones such as the DIFC and ADGM have their own data protection regimes as well.

Global routing in cloud model APIs

Many cloud model APIs can route a request to whichever region has capacity. Some deployment options from major providers state plainly that prompts and responses may be processed in any geography where the model runs, even when stored data stays in the region you selected. Region-pinned options exist, but they do not always offer the newest models, and the bank still has to verify the routing behavior for every model and deployment type it uses.

None of this means a UAE bank can never use an external model. It does mean that for workloads touching customer data, keeping inference in the country is often the shortest path to approval. This article is not legal advice, and the specifics for your institution and data types should be confirmed with your compliance and legal teams.

Which open-weight models can a UAE bank run privately?

Open-weight models are the foundation of a private deployment because you can download the weights and run them anywhere. The UAE has produced several of its own.

  • Falcon from the Technology Innovation Institute (TII) in Abu Dhabi. The Falcon family spans several sizes and generations, and recent releases include variants aimed at Arabic.
  • Jais, developed in the UAE for Arabic and English. It is a natural candidate for customer-facing text, Arabic correspondence and documents that mix both languages.
  • K2 Think from MBZUAI and G42, a 32-billion-parameter reasoning model built on Qwen 2.5. Its size makes it practical to serve on a single node.
  • DeepSeek, Llama and Qwen families from outside the region. These are widely used, well supported by serving software, and often the strongest options for English-heavy analytical work.

Pick the model by testing it on your own documents and tasks. Public leaderboards are a starting point at most. A model that scores well on math benchmarks may still mishandle a scanned Arabic trade-finance document. Licensing terms also differ between families and between versions, so legal review of the license belongs in the selection step.

A reference architecture for a private LLM in a bank

A production deployment has more parts than the model. A minimal version that satisfies most security reviews looks like this.

In-country GPU capacity

Inference runs on GPUs in the bank's own data center or in a UAE facility under contract. The number and type of GPUs determine which models you can serve and at what throughput, so sizing comes early. A large mixture-of-experts model may need a full multi-GPU node, while a 7–32B model can often run on one or two cards.

The serving stack

An inference server loads the model, batches incoming requests and streams tokens back. This layer decides most of your latency and cost. It is where quantization, parallelism across GPUs and scheduling are configured, and where custom kernels plug in if the defaults are too slow.

Retrieval over internal documents

Most banking use cases depend on the bank's own corpus: policies, product terms, credit files, regulatory circulars. A retrieval layer indexes these documents inside the perimeter and passes the relevant passages to the model with each request. Every answer can then cite the source it came from, which is what reviewers ask for first.

Identity and access control

The LLM service should sit behind the bank's existing identity provider. Retrieval must respect document-level permissions, so a user never receives a passage from a file they could not open directly. This is the step most pilots skip and most production reviews catch.

Logging for audit

Each request should be logged with the user, the prompt, the retrieved documents, the output and the model version that produced it, all stored inside the perimeter. With that record, a compliance officer or internal auditor can reconstruct what the system did and why, months later.

Performance and cost: why throughput decides on-prem economics

With an API you pay per token. On-prem you pay for GPUs whether they are busy or idle, so the cost of each request depends on how many requests the hardware can serve per second. Doubling throughput roughly halves the cost per request on the same hardware. That makes inference engineering the main lever on the business case, more than the hardware discount.

The usual techniques:

  • Quantization. Storing weights in 8-bit or 4-bit formats cuts memory and speeds up generation, often with little quality loss. "Often" is the important word: quality must be measured on your own evaluation set before a quantized model ships.
  • Parallelism. Large models are split across GPUs with tensor or pipeline parallelism. The right split depends on the model, the interconnect and the traffic pattern.
  • Batching and scheduling. Serving many requests together keeps GPUs busy. Good scheduling raises throughput while keeping latency within the target.
  • Custom kernels. When profiling shows a specific operation dominating, a hand-written CUDA kernel can recover performance that general-purpose libraries leave on the table.

Our DeepSeek R1 FP4 inference world record, 303 tokens per second on DeepSeek R1 in FP4 on NVIDIA Blackwell, came from FP4 inference, speculative decoding and a custom TensorRT-LLM build, developed with Avian.io and NVIDIA. On a bank's own hardware, the same techniques change how many GPUs a workload needs and how fast each user gets an answer. In our financial deployments, the same engineering keeps P99 inference latency under 50 ms. More on this work is on our LLM engineering page.

Where should a bank start with a private LLM?

Start with one document-heavy workflow. Good candidates share a few traits: a team spends hours reading, extracting or reconciling documents; the documents already live inside the perimeter; and success can be measured in review time or error rate.

Examples include onboarding and KYC packs, credit memo preparation, trade-finance document checks and internal policy questions. A first project like this exercises every part of the architecture (GPUs, serving, retrieval, access control, audit logging) without putting the model in front of customers.

A sensible sequence:

  1. Map the workflow. Document the inputs, the decisions, the controls and the metric that defines success.
  2. Benchmark candidate models on a sample of real documents, including Arabic ones if they appear in the workflow.
  3. Size the hardware from the measured throughput and the expected volume.
  4. Build the pipeline with retrieval, access control and logging in place from day one.
  5. Run it beside the existing process, compare outputs with reviewer decisions, then expand.

Our team has built on-prem LLM tooling and document intelligence inside a tier-1 banking perimeter at JP Morgan, and our banking document intelligence work delivered a 60% reduction in manual review. The constraints in a UAE bank are similar: the model is half the problem, and the controls around it are the other half. Our banking and finance page covers how we approach these engagements.

Talk to us about a private LLM deployment

Archeon has offices in San Francisco and Dubai, and the team brings operator experience from NVIDIA, Siemens Healthineers and JP Morgan. If you are scoping a sovereign AI deployment in the UAE and want a second opinion on models, hardware or architecture, get in touch.

Related projects

Related capabilities

More articles