Engineering7 min read

Inside the DeepSeek R1 World Record: 303 Tokens per Second in FP4 on NVIDIA Blackwell

An engineering explainer on fast DeepSeek R1 inference: why a 671B Mixture-of-Experts reasoning model is hard to serve quickly, and how FP4 inference, speculative decoding and a custom TensorRT-LLM build on a single NVIDIA DGX Blackwell node fit together behind the 303 tokens per second world record we set with Avian.io and NVIDIA.

DeepSeek R1 is one of the most capable open-weight reasoning models available, and one of the most demanding to serve fast. In April 2025, after about a month of work with Avian.io and NVIDIA, our team's inference stack reached 303 output tokens per second on DeepSeek R1 in FP4, on a single NVIDIA DGX Blackwell node, using a custom version of TensorRT-LLM (TRT-LLM). Artificial Analysis benchmarked the Avian.io private API endpoint independently and called it "the fastest speed we have measured yet for DeepSeek R1", a world record at the time of their post.

The project page summarizes the outcome. This article explains the engineering context: where the time goes when you serve a model like R1, and how FP4, model-specific kernels and speculative decoding combine. We describe how each technique works, so you can judge where it applies to your own models.

Why does decode speed matter so much for R1?

R1 is a reasoning model. Before it gives an answer, it writes out a chain of thought that often runs to thousands of tokens. A user sees nothing useful until much of that reasoning is done.

That makes per-user decode speed, the rate at which one request produces tokens, the number that decides whether the model feels usable. A 3,000-token reasoning trace takes 100 seconds at 30 tokens per second and about 10 seconds at 300. For interactive assistants and for agents that chain several reasoning calls, that difference decides whether a product works at all.

Why is DeepSeek R1 hard to serve fast?

R1 shares its architecture with DeepSeek V3, the base model it was trained from. Three properties of that architecture shape every serving decision.

671B parameters, about 37B active per token

R1 is a Mixture-of-Experts (MoE) model with 671B total parameters, of which about 37B are activated for each token. Each MoE layer holds a large pool of routed experts plus a shared expert, and a router picks a small subset of routed experts for every token. Compute per token looks like a ~37B dense model. Memory looks like a 671B model, because every expert has to be resident somewhere in case the router selects it.

The official weights ship in FP8, so the model needs roughly 671 GB for weights alone, before any KV cache or activations.

Multi-head Latent Attention

The model uses Multi-head Latent Attention (MLA), which compresses keys and values into a low-rank latent vector per token instead of storing full per-head keys and values. That shrinks the KV cache and makes long reasoning traces more affordable. It also means attention kernels written for standard multi-head or grouped-query attention do not map cleanly onto it. MLA needs its own kernels to run at full speed.

Decode is memory-bandwidth bound

During decode, each step produces one token per sequence. For every step, the GPU must read the weights it needs from memory, do a relatively small amount of arithmetic on them, and move on. When the goal is speed for a single request, batches are small, arithmetic intensity is low, and the tensor cores wait for bytes to arrive. Decode speed is set by memory bandwidth and by the fixed overheads of each step.

FP4 inference: fewer bytes per token

If decode speed is set by bytes moved per token, the most direct lever is to make each weight smaller. FP4 stores values as 4-bit floating-point numbers. Compared with the FP8 release, that halves the bytes read per weight, and it frees memory for KV cache.

FP4 differs from INT4 in how the 16 available levels are spread. A floating-point format places more levels near zero, where most weights sit, and uses shared scale factors per small block of values to handle the range. NVIDIA's Blackwell generation, for example, supports a 4-bit floating-point format (NVFP4) in hardware, with a shared scale for each block of 16 values, so tensor cores can consume it directly. On GPUs without native FP4 support, the format falls back to slower paths or is not supported at all, so hardware and format have to be chosen together. Our record ran on Blackwell, which supports NVFP4 natively.

The hard part is accuracy. Four bits leave very little room, and reasoning models can be sensitive: a small error early in a long chain of thought can carry through to the final answer. Techniques commonly used to manage this include:

  • Fine-grained block scales, so outlier values only affect their immediate neighbors.
  • Calibration on representative data, so the quantization minimizes error on the activations the model actually sees.
  • Selective precision, keeping layers that prove sensitive at higher precision. In an MoE model the expert weights make up the bulk of the parameters, so quantizing them captures most of the savings.

For the record run, Artificial Analysis reported that the FP4 version maintained accuracy across their evaluation suite compared with the native FP8 version. We cover how to measure the quality cost of 4-bit formats in our guide to INT4 and FP4 quantization quality loss.

Kernel-level optimization

A quantization format is only as fast as the kernel that runs it. For the record, we built on a custom version of NVIDIA's TensorRT-LLM. General-purpose kernels in common libraries are written to be correct and reasonably fast across many shapes, data types and models. A kernel written for one model's exact matrix shapes, one number format and one hardware target can make choices a general kernel cannot.

Kernel work of this kind usually targets three things.

Fusion. A naive forward pass launches a separate kernel for each small operation: dequantize, matrix multiply, apply activation, scale by the router weight, and so on. Each launch writes its output to memory and the next one reads it back. Fusing these steps keeps intermediate values in registers and shared memory. On a bandwidth-bound workload, every avoided round trip is time recovered.

Memory traffic. Decode speed tracks the number of bytes moved per token, so kernels for this kind of model are built around that number: tiling so weights are read once per step, access patterns that keep loads coalesced, and grouping tokens by expert so each expert's weights are streamed through as few times as possible.

Per-step overhead. At small batch sizes, the gaps between kernels matter. Launch latency, synchronization and host-side scheduling can take a noticeable share of each step. Reducing the number of launches and capturing the decode loop so it replays with little host involvement keeps the GPU working instead of waiting.

Speculative decoding: more than one token per step

Speculative decoding was the third part of the record setup. Even with fast kernels, standard decoding produces one token per forward pass of the full model. Speculative decoding changes that ratio. A lightweight draft, either a small separate model or extra prediction heads on the main model, proposes several tokens ahead. The full model then checks all of them in a single forward pass and keeps the longest run it agrees with.

When drafts are accepted, the model emits several tokens for roughly the cost of one step. When they are rejected, the model falls back to its own prediction, so with standard verification rules the output matches what the full model would have produced on its own. DeepSeek's own architecture includes a multi-token prediction module, which is one natural source of drafts for this family of models.

The gain depends on the acceptance rate, which depends on the text. Long reasoning traces contain a lot of predictable structure, which tends to suit speculative decoding. The cost is extra compute per step for the draft and for verification, which a bandwidth-bound decode loop usually has to spare.

How the techniques fit together

These techniques build on each other. FP4 cuts the bytes each step must read. Kernels written for the format and the hardware make sure that saving reaches the GPU instead of being lost to dequantization and overhead. Speculative decoding then makes each of those faster passes yield more than one token.

Optimize any one in isolation and the bottleneck moves somewhere else. Fast kernels at FP8 still read twice the bytes. FP4 through general-purpose kernels can lose much of its advantage. Speculative decoding on a slow verification pass multiplies a small number.

How to check that a faster model is still the same model

A speed number without a quality number is not useful, and this matters more for a reasoning model than for most. Some practices we consider non-negotiable for any low-precision work:

  • Compare against the higher-precision reference on the same harness. Same prompts, same decoding settings, same scoring.
  • Use more than perplexity. Perplexity is a quick sanity check, but it can hide regressions on reasoning, math, code or long-context tasks.
  • Test the tasks that matter to the deployment. Aggregate benchmark scores can mask a drop on the one capability a product depends on.
  • Watch long generations. Small numerical errors can accumulate over thousands of decoded tokens, which is exactly what reasoning models produce.
  • Account for noise. Run enough samples, and repeat with different seeds where sampling is involved, to know whether a measured difference is real.

What this means for teams serving large models on their own GPUs

Faster decode per user changes which products are possible. A reasoning model that answers in seconds can sit inside an interactive workflow; one that takes minutes cannot. Smaller weights also change the hardware calculation, since a model that fits in fewer GPUs is cheaper to run and simpler to operate.

That matters for on-prem and in-country deployments, where organizations run open-weight models inside their own perimeter for data residency, regulatory or latency reasons. The work is specific to a model, a hardware generation and a traffic pattern. Profiling real traffic comes before any optimization, and quality evaluation comes before any change of precision. Our LLM engineering engagements follow that order.

How do vLLM and SGLang fit in?

Open-source serving engines such as vLLM and SGLang handle much of what production serving needs: continuous batching, paged KV cache management, request scheduling, multi-GPU execution, and support for many quantization formats. Both have added DeepSeek-specific optimizations, including MLA support. For most teams they are the right starting point.

Model-specific kernel work sits at a different level. It targets the remaining headroom for one model on one class of hardware, the kind of gap that general-purpose engines, built to support many models, are not designed to close for every case. Teams that need the fastest possible responses from a single flagship model are the ones for whom that extra work pays off.

Working with us

We do this kind of work for AI labs and infrastructure teams in frontier compute and for organizations deploying private LLMs on their own hardware. If you are serving a large model and want to know how much speed your GPUs are leaving unused, get in touch.

Related projects

Related capabilities

More articles