Capability

Sovereign AI, On-Prem LLMs & Inference Optimization

On-prem and air-gapped LLM deployment with custom kernels, low-precision inference and speculative decoding, the methods behind our DeepSeek R1 inference world record on NVIDIA Blackwell with Avian.io and NVIDIA.

Large language models are easy to demo and hard to operate. The gap shows up as latency that will not fit a workflow, GPU bills that grow faster than usage, or a security review that rules out any external API. Our LLM practice exists to close that gap: we make models run inside your perimeter, fast enough and cheaply enough to be used every day.

Sovereign AI and data residency

More organizations now need AI that runs entirely inside their own jurisdiction. In the UAE, for example, sector rules from the Central Bank and health data law require most regulated data to stay in the country, and many cloud model APIs route requests to data centers abroad. Sovereign AI solves this by keeping the model, the data and the hardware under your control. We deploy private LLMs on your own GPUs or in an in-country data center and engineer them to be fast and cheap enough for daily use.

What we build

  • On-prem and air-gapped inference stacks that run open-weight models on your hardware, with no third-party model APIs in the request path.
  • LLM inference optimization: custom CUDA kernels, FP4 and int4 quantization, speculative decoding and tensor parallelism across multi-GPU nodes.
  • Document intelligence and retrieval systems that ground model outputs in your own corpus and keep a record of what was retrieved and why.
  • Evaluation harnesses that measure quality on your tasks, so every optimization is checked against a number that matters.
  • Serving infrastructure with the logging and versioning needed for audit and rollback.

Where we have shipped it

Our clearest public result is the DeepSeek R1 FP4 Inference World Record: with Avian.io and NVIDIA, our team built the inference stack that reached 303 output tokens per second on DeepSeek R1 in FP4, on a single NVIDIA DGX Blackwell node, combining FP4, speculative decoding and a custom TensorRT-LLM build. Artificial Analysis benchmarked it independently in April 2025 and called it the fastest DeepSeek R1 speed they had measured.

For Believe, we built a multi-agent contract-review pipeline on Gemini and LangGraph with schema-constrained generation, taking contract type classification from 40% to 98.2% and field-level extraction to 91-93% concordance against reviewer ground truth. See the Believe contract review case study.

The same team has built on-prem LLM tooling and document intelligence inside a tier-1 banking perimeter. That experience shapes how we approach Banking & Finance engagements, where the model is only half the problem and auditability is the other half.

How an engagement runs

Each LLM project runs in four phases, with your engineers involved from the start.

  1. Scoping. We pin down the workload: request volume, context lengths, latency targets, hardware and data constraints. We agree on the quality metrics that any optimization must preserve.
  2. Research. We benchmark candidate models and serving configurations on your tasks, then profile to find the real bottleneck, whether that is memory bandwidth, scheduling or a specific kernel.
  3. Production. We implement the changes, integrate with your identity, logging and storage, and load-test against realistic traffic.
  4. Handover. Your engineers get the code, configurations, benchmarks and runbooks, and we pair with them until they can upgrade models and re-tune the stack themselves.

Deployment and stack considerations

  • Hardware first. Throughput and cost depend on the GPUs you own or can procure. We design the parallelism and quantization strategy around that hardware, not the other way around.
  • Quality is measured, not assumed. Every quantization or decoding change is validated against your evaluation set before it ships.
  • Perimeter and audit. For regulated environments we keep inputs, outputs and model versions logged inside your infrastructure, so a reviewer can reconstruct what the system did.
  • Ownership. We favor open-weight models and open-source components where they serve the work, so you are not locked into a vendor to keep the system running.

For compute-bound problems beyond LLM serving, such as training infrastructure, see R&D & Frontier Compute. To discuss your workload, get in touch.

Frequently asked questions

What is sovereign AI?

Sovereign AI means the model, the data it processes and the hardware it runs on all stay under your control and inside your jurisdiction. In practice that means open-weight models deployed on infrastructure you own or rent in-country, with no prompts or documents sent to an external model API.

Why run an LLM on-prem instead of calling a model API?

For regulated organizations, sending documents to a third-party API is often not an option. On-prem and air-gapped deployments keep data, prompts and outputs inside your perimeter, and give you control over model versions, latency and cost.

How much quality do you lose with quantization?

It depends on the model and the method, which is why we measure it on your tasks rather than assume it. Our DeepSeek R1 world record ran in FP4, with accuracy held against the native FP8 version per Artificial Analysis, and any change of precision we ship is checked on an evaluation set built from your tasks first.

Can you optimize a model we already deploy?

Yes. We profile the existing serving stack, find where time and memory go, and work down from there: batching and scheduling, quantization, parallelism strategy and, where it pays off, custom kernels.

Which open-weight models do you deploy?

We pick the model by benchmarking candidates on your tasks, not by default. Candidates usually include open-weight families such as DeepSeek, Llama and Qwen, and Falcon or Jais for Arabic-language workloads.

Do you fine-tune models as well?

Where the task calls for it, yes. Many engagements are as much about evaluation, retrieval and serving as about the weights, so we decide with you whether fine-tuning is worth its cost before starting it.

Selected work

Industries we apply it in

Have a llm engineering problem worth solving properly?

Book an intro call