产品多源确认精选

PyTorch 的 Helion 内核 DSL 集成 vLLM,部分工作负载吞吐提升超 10%

精选理由

PyTorch 官方把 Helion 接进 vLLM,一套 GEMM 代码自动选最优算法,在 Hopper 上跑赢 CUTLASS 和 DeepGEMM,搞推理部署的可以看看。

Helion 是 PyTorch 团队推出的内核 DSL,官方博客展示了如何将其集成到 vLLM 的线性后端。一套 Helion GEMM 实现覆盖 Standard GEMM、Split-K 和 Swap-AB 三种算法变体,按张量形状自动选择最优配置。在 NVIDIA Hopper GPU 上,结合逐形状调优与混合调度,性能超过 vLLM 默认的 CUTLASS 和 DeepGEMM 后端,部分工作负载端到端吞吐提升超过 10%。文章作者来自 Red Hat 和 Meta 的 PyTorch 团队。

原文 · PyTorch

Helion, PyTorch's kernel DSL, is showing what autotuned high-level kernels can do for production inference.

In this post, we integrate Helion into @vllm_project linear backend and show how a single Helion GEMM implementation can cover multiple algorithmic variants — Standard GEMM, Split-K, and Swap-AB — with the best variant and config selected automatically per shape. On @nvidia Hopper GPUs, combining per-shape tuning with hybrid dispatch outperforms vLLM's default CUTLASS and DeepGEMM backends across the evaluated models, with consistent end-to-end gains and more than 10% throughput improvement on some workloads.

Read our latest blog: https://t.co/34ni2qmEes

✍️ Sean Chen (@RedHat) and Shangdi Yu (PyTorch, @Meta Platforms)