Engineering
2 articles · all posts
Inside the DeepSeek R1 World Record: 303 Tokens per Second in FP4 on NVIDIA Blackwell
An engineering explainer on fast DeepSeek R1 inference: why a 671B Mixture-of-Experts reasoning model is hard to serve quickly, and how FP4 inference, speculative decoding and a custom TensorRT-LLM build on a single NVIDIA DGX Blackwell node fit together behind the 303 tokens per second world record we set with Avian.io and NVIDIA.
INT4 Quantization for LLMs: How to Measure the Quality Loss Before You Ship
A practical guide to INT4 and FP4 LLM quantization: how GPTQ, AWQ, FP8, NVFP4 and KV cache quantization work, why perplexity hides regressions, and how to build an evaluation plan that tells you whether a quantized model is safe to ship.