LLM Optimization

DeepSeek V3 Inference World Record

Achieved world-record throughput on DeepSeek V3 through custom CUDA kernels, int4 quantization with minimal quality loss, and a novel tensor parallelism strategy across 8-GPU clusters.

DeepSeek V3 Inference World Record
World record3.2× speedup< 1% quality loss