GPU Performance Engineering for AI Labs
GPU performance engineering for teams limited by throughput or cost: custom CUDA kernels, quantization, distributed inference and training infrastructure. We contribute to open source where it helps the work.
- DeepSeek R1 · FP4
- 303 tok/s
- LLM inference speed
- World record
- Precision
- FP4
For AI labs and compute-heavy teams, the limit is often not ideas but hardware. Training runs queue for weeks, serving costs grow faster than usage, and new models arrive faster than the infrastructure can adapt. Framework defaults leave throughput unused, and the fixes live in kernels, memory layouts and communication patterns that few teams have time to master. That is the layer we work at.
The problems we see in compute-bound teams
- Serving costs that do not scale. A model that is affordable for a pilot becomes the largest line item at production volume.
- Latency that rules out use cases. Long contexts and large models push response times past what an interactive product can tolerate.
- Underused GPUs. Profiling often shows hardware waiting on memory, communication or poorly fused operations rather than doing math.
- Training infrastructure that does not compound. Every experiment re-solves the same scaling and reliability problems instead of building on a shared base.
What we build
- Custom CUDA kernels for the operations that dominate your profile, including CUDA and TensorRT pipelines.
- Quantization to FP4, int4 and other low-precision formats, validated on your quality metrics before it ships.
- Distributed inference: tensor parallelism and scheduling strategies for multi-GPU and multi-node serving.
- Training infrastructure: distributed training and accelerated pipelines that make each experiment cheaper and faster than the last.
The engineering practice behind this is described in LLM Engineering, and open-ended performance research runs through Research & Development.
Constraints we design for
Quality is a hard constraint
Speed that degrades outputs is not a gain. We measure quality on your evaluations at every step, before and after every change of precision.
Your hardware, your stack
We optimize for the GPUs you have or can procure and integrate with your existing serving and training systems, rather than requiring a platform migration.
Reproducibility
Benchmarks are scripted, versioned and repeatable, so improvements can be verified by your team and defended to stakeholders.
Relevant work
Our DeepSeek R1 FP4 Inference World Record is the clearest example of this work: a world record we set with Avian.io and NVIDIA, reaching 303 output tokens per second on DeepSeek R1 in FP4 on a single NVIDIA DGX Blackwell node, verified by Artificial Analysis. Our team also brings operator experience from NVIDIA in LLM inference optimization, CUDA and TensorRT pipelines and distributed training infrastructure.
For Wearit, we trained a virtual try-on diffusion model from scratch in distributed multi-GPU runs (A100/H100) on the Jean Zay national supercomputer under a GENCI compute grant.
How to start
We usually begin with a profiling engagement: you share the workload, hardware and quality bar, and we return a measured breakdown of where time and memory go and which changes are worth making. From there we implement, benchmark and hand over the code and runbooks to your engineers. Get in touch to discuss your workload.
Selected partners: NVIDIA · Avian
Frequently asked questions
Who is this for?
AI labs, model providers and engineering teams whose roadmap is limited by GPU throughput, memory or cost. If your models work but are too slow or too expensive to serve or train at the scale you need, this is the problem we work on.
What did the DeepSeek R1 world record involve?
With Avian.io and NVIDIA, our team built the inference stack that reached 303 output tokens per second on DeepSeek R1 in FP4, on a single NVIDIA DGX Blackwell node, using FP4, speculative decoding and a custom TensorRT-LLM build. Artificial Analysis benchmarked it independently in April 2025 and reported that accuracy held against the native FP8 version.
Do you only work on inference?
No. We also build distributed training infrastructure and accelerated training pipelines. Inference is where the gains are often most visible, but training efficiency compounds across every experiment.
Will you contribute changes upstream?
Where it serves the work and you agree, yes. We make open-source contributions when they help the system, and keep proprietary work private when it does not.
Selected work
- LLM OptimizationDeepSeek R1 Inference World Record on NVIDIA BlackwellWith Avian.io and NVIDIA, our team built the inference stack behind a DeepSeek R1 world record: 303 output tokens per second in FP4 on a single NVIDIA DGX Blackwell node, independently benchmarked by Artificial Analysis.
- Generative AIVirtual Try-On Diffusion Model for WearitArchitected Wearit's core generative AI pipeline for virtual clothing try-on, from R&D to real-time production: a diffusion model built from scratch, trained on the Jean Zay supercomputer on 3M+ labeled fashion images, serving results in under 2 seconds.
Capabilities we bring
- CapabilityLLM EngineeringOn-prem and air-gapped LLM deployment with custom kernels, low-precision inference and speculative decoding, the methods behind our DeepSeek R1 inference world record on NVIDIA Blackwell with Avian.io and NVIDIA.
- CapabilityResearch & DevelopmentApplied AI research for problems without a known solution. We co-author with UC Berkeley, Harvard and UCL, then deploy what works into your stack.
- CapabilityAI AgentsAgentic workflows that use your tools and data, run on models inside your perimeter, and keep a human approval step and an audit log for every action.
Have a r&d & frontier compute problem worth solving properly?
Book an intro call