AI for Banks & Asset Managers
We build risk engines, on-prem LLM stacks and document-intelligence systems for tier-1 banks and asset managers. Auditable from training data to inference and deployed inside your perimeter, with no third-party model APIs.
- Manual review reduction
- 60%
- P99 inference latency
- < 50ms
- Data leaving perimeter
- 0
Banks and asset managers do not lack AI ideas. What they lack is a path from pilot to production that a CISO, a model risk committee and an auditor will all sign off on. Public model APIs are often ruled out on day one, internal platforms move slowly, and many vendors cannot explain how their models reach a decision. We work inside those constraints rather than around them.
The problems we see in financial institutions
- Review-heavy operations. Teams read, extract and reconcile large volumes of documents by hand, from onboarding packs to contracts and filings. Throughput is capped by headcount, and quality varies by reviewer.
- Generative AI that cannot leave the building. Analysts want LLM tooling, but security policy prohibits sending client or position data to external providers.
- Models that stall in validation. Risk and research models are built quickly but get stuck in review because lineage, reproducibility and limitations are not documented.
- Latency and cost at scale. On-prem models that work in a pilot become too slow or too expensive once real volumes arrive.
What we build for banks and asset managers
- Document-intelligence systems that extract, classify and summarize documents for review teams, with every output traceable to its source. Our work in this sector has delivered a 60% reduction in manual review.
- On-prem LLM stacks for analysts and operations, running open-weight models on your hardware with retrieval over your own corpus.
- Risk engines and compliance models with versioned data, documented assumptions and drift monitoring.
- Research infrastructure for alpha and risk teams that keeps experiments reproducible and reviewable.
The capability behind this work is described in Finance AI and LLM Engineering.
Constraints we design for
Regulation and audit
Every system is auditable from training data to inference. We version datasets, code and weights, log each prediction with the model version that produced it, and write documentation for validators and auditors.
The perimeter
We deploy inside your infrastructure with no third-party model APIs. Zero data leaves the perimeter. Identity, access control and logging integrate with what your security team already runs.
Latency
Trading-adjacent and customer-facing workflows cannot wait on a slow model. We engineer serving stacks to P99 inference latency under 50 ms, using the same kernel and low-precision techniques behind our DeepSeek R1 inference world record.
Operator experience
Our team has built on-prem inference stacks, risk and compliance modeling and audit-ready ML systems inside a tier-1 banking perimeter at JP Morgan, and has worked with Avian. We know how procurement, security review and model validation work in a large institution, and we plan the engagement around them from the start.
Relevant work
- Identity Verification Engine: the AI engine of Oneex's identity-verification platform, used in banking among other critical environments, with MRZ OCR at 99.65% character accuracy, multi-spectral forgery detection and on-device biometrics.
- Industrial OCR Engine Audit: benchmarking and architecture audit for Tessi, a European leader in document processing.
- DeepSeek R1 Inference World Record on NVIDIA Blackwell: 303 tokens per second on DeepSeek R1 in FP4, verified by Artificial Analysis, the kernel and low-precision work behind our on-prem LLM serving.
How to start
We usually begin with a short scoping phase on one workflow: we map the process, the data, the controls it must satisfy and the metric that defines success. From there we build, validate and deploy inside your environment, then hand over the system, documentation and runbooks to your team. Book an intro call to discuss a specific workflow.
Selected partners: JP Morgan · Avian
Frequently asked questions
Can generative AI be used in a bank without sending data to a vendor?
Yes. We deploy open-weight models on your own infrastructure, with no third-party model APIs in the request path. Documents, prompts and outputs stay inside your perimeter and under your existing security controls.
Can a bank in the UAE use LLMs and keep customer data in the country?
Yes. We deploy open-weight models on GPUs inside your data center or an in-country facility, so prompts, documents and outputs never leave the jurisdiction. Logging and access control integrate with your existing systems so compliance teams can audit every request.
How do you support model risk management?
We build systems that are auditable from training data to inference: versioned data and models, reproducible training, logged predictions and documentation written for validators. That gives your model risk team the evidence it needs to review the system.
Is on-prem inference fast enough for production workloads?
It can be, when the serving stack is engineered for it. We apply kernel, quantization and parallelism work from our inference research, and our financial deployments run at P99 inference latency under 50 ms.
Where does a typical engagement start?
Usually with one high-volume, review-heavy workflow where the cost of manual work is clear and the data is already inside your perimeter. Document intelligence for review teams is a common first project.
Selected work
- Computer VisionAI Identity Document Verification Engine for OneexDesigned and built the AI engine of Oneex's identity-verification platform, which processes over 10 million identity documents a year in airports, banking, defence and sovereign infrastructure. OCR, multi-spectral authenticity analysis, print-technique forensics and biometrics, all running on CPU-only edge hardware.
- LLM OptimizationDeepSeek R1 Inference World Record on NVIDIA BlackwellWith Avian.io and NVIDIA, our team built the inference stack behind a DeepSeek R1 world record: 303 output tokens per second in FP4 on a single NVIDIA DGX Blackwell node, independently benchmarked by Artificial Analysis.
- Document AIIndustrial OCR Engine Audit & Roadmap for TessiAdvised the Technical Direction of Tessi, a European leader in document processing, on the evolution of their industrial OCR engine: a benchmark against state-of-the-art models, an end-to-end architecture audit and a technical roadmap for the next-generation platform.
Capabilities we bring
- CapabilityFinance AIRisk models, alpha research tools and document intelligence for banks and asset managers, built to pass model validation, audit and stress testing.
- CapabilityLLM EngineeringOn-prem and air-gapped LLM deployment with custom kernels, low-precision inference and speculative decoding, the methods behind our DeepSeek R1 inference world record on NVIDIA Blackwell with Avian.io and NVIDIA.
- CapabilityAI AgentsAgentic workflows that use your tools and data, run on models inside your perimeter, and keep a human approval step and an audit log for every action.
Have a banking & finance problem worth solving properly?
Book an intro call